Skip to content

feat: input artifacts in server and CLI task requests - #844

Merged
0xShug0 merged 4 commits into
0xShug0:mainfrom
Liquid4All:request-artifacts
Oct 9, 2026
Merged

0xShug0 merged 4 commits into
0xShug0:mainfrom
Liquid4All:request-artifacts

Conversation

@ykhrustalev

Copy link
Copy Markdown
Contributor

Server and CLI task requests can carry input artifacts, in the shape results already return them in. With lfm2_audio's multi-turn S2S (#838) on main, this lets server and CLI clients hold a conversation.

Problem

  • Results return artifacts in an artifacts array, but no request route reads one. Only the C API fills TaskRequest::input_artifacts, so server and CLI clients cannot send back what a result gave them
  • lfm2_audio S2S (feat(lfm2_audio): multi-turn speech-to-speech with client-carried history #838) takes a conversation's earlier turns as input artifacts, lfm2_audio.question and lfm2_audio.reply, and with return_codes it returns each turn's lfm2_audio.reply. The server returns that reply but cannot take it back. So through the server and the CLI, every S2S request is a first turn, and the lfm2_audio docs point conversations to the C API

What this PR changes

  • build_request_from_json reads an artifacts array, in order, into TaskRequest::input_artifacts. That covers /v1/tasks/run, /v1/tasks/stream and each /v1/tasks/batch entry on both server runtimes, plus the CLI's --request-sequence JSON and workflow requests
  • An entry has the shape the server writes: a non-empty id (ids may repeat), a kind as the server names it, a Base64 payload or data URI, and optional meta. Numbers and booleans in meta become text, as options values do (1.0 becomes 1). path may replace payload, and it resolves the way the request's audio does: on the server against its working directory, in the CLI against the JSON file
  • A malformed entry throws engine::runtime::InvalidRequestError, which main already has, with a message naming the entry, such as artifacts[1] (example.tokens): unknown kind 'tokens'. Both server runtimes answer it with 400 invalid_request_error, and the CLI exits 1
  • AuK, HeartMuLa, VibeVoice, VoxCPM1, VoxCPM2 and YuE2 refuse any input artifact, as they do through the C API, and now throw that error to do it. Their other refusals are unchanged
  • max_request_body_bytes bounds inline payloads, as it bounds audio_base64. One request's payloads, inline and from files, may total 2 GiB, the default body limit. A static_assert keeps the two equal
  • A static_assert checks that the kind table names every ArtifactKind in enum order, up to Custom, which stays last. Inline payloads decode straight into the artifact (new base64_decode_bytes)
  • The first commit moves the Base64 codec, unchanged, from app/server to app/common (namespace minitts::app) so the CLI can use it. server_base64_test becomes base64_test
  • Docs: the server README, docs/usage.md and docs/c_api.md. The last commit updates the lfm2_audio page, which said the server takes no request artifacts. Its Conversations section now shows a later turn through /v1/tasks/run, with the earlier question as a path and the reply's payload and meta as the result returned them. It also says the CLI's request JSON takes the same array, and that the CLI's payload_hex results go back as a file's path. The limitation now names only the live route. The server README now says what a relative path resolves against

What changes for users, and what does not

  • Requests can send artifacts, and which ones a model reads is up to its family. lfm2_audio S2S reads a conversation's history, so server and CLI clients can now carry one: for each earlier turn, in order, the question's WAV bytes as lfm2_audio.question (kind custom), then the lfm2_audio.reply that turn returned. The rules are feat(lfm2_audio): multi-turn speech-to-speech with client-carried history #838's: S2S turns away other ids and a history out of order, while ASR and TTS turn away lfm2_audio.* ids and ignore the rest
  • The six families above refuse any artifact, and every family other than them and lfm2_audio ignores them. Requests without artifacts are unchanged
  • The live route, /v1/audio/speech/live, takes no artifacts, so each of its requests is still a first turn
  • A malformed artifacts key used to be ignored and now gets a 400. So does a non-empty one sent to the six families, which used to run (the CLI exits 1 for both). So does a history that lfm2_audio turns away; before, the server dropped the key and answered a first turn
  • Still 500 server_error, as before: a body over max_request_body_bytes, and request errors that throw a plain std::runtime_error, such as an unknown option or an option value the shared parsers reject. In this PR, only the request reader and the six refusals throw InvalidRequestError
  • path reads any regular file the server process can read, as audio and voice_ref do. No model echoes its input artifacts. If an lfm2_audio.reply's path names some other file, the refusal message can quote an out-of-range 32-bit value from that file, or its size
  • The 2 GiB total does not follow a lower max_request_body_bytes. Each batch entry is one request, and a batch reads all its entries before it runs, so with path artifacts a batch can hold 2 GiB per entry
  • The CLI writes result artifacts as JSON with payload_hex, which requests do not take. Write the bytes to a file, or Base64 them

Testing
At this head, on an A10 host: its CPU (a Xeon 8358, 8 threads) and its A10 (CUDA). Release builds with gcc 12, with lfm2_audio and the six families built in. Main still has a ggml-cuda race that can change a reply at a near-tie code (#837 fixes it), so the CUDA runs used a build with #837's fix merged in. Metal ran only at an earlier revision (below).

  • Build: the warnings are the same as main's at ef860c9, and none is in a changed file
  • ctest, filtered to the server, HTTP, CLI, Base64, C API, framework, lfm2_audio and YuE2 request tests: 42 of 42 here (40 passed, 2 skipped) and 41 of 41 on main at ef860c9 (39 passed, the same 2 skipped), lfm2_audio_encoder_cpu_repack_test included. The skipped two are the C API model tests, which need a model root. The two sets differ only by the new yue2_request_test and the base64_test rename. Three more runs of the same filtered set on each tree passed, parallel_http_live_body_test included. The PR's unit tests cover every payload form and kind name, the rejections other than file I/O failures, the 2 GiB file total (skipped on Windows), run, stream and batch end to end on the parallel runtime, and the legacy runtime's 400s
  • Conversations on the CPU, EN and JP Q4_0: the C API ran an EN conversation of 10 turns, offline and streamed, and a JP one of 3 turns offline, with return_codes and a fixed seed per turn. Each server and CLI turn sent the C API's earlier questions and replies as history. Each one returned the C API's reply artifact (payload, kind and meta), text and audio byte for byte, 384 of 384 checks:
    • Legacy runtime: EN /v1/tasks/run (questions as path, replies Base64) and /v1/tasks/stream (both Base64, also checking the events' audio and partial text), 10 turns each
    • Parallel runtime (--parallel-jobs, one slot): EN /v1/tasks/run (both Base64) and /v1/tasks/stream (both path), 10 turns each, and JP /v1/tasks/run, 3 turns
    • meta sent back as JSON numbers and booleans, 3 turns per runtime: same replies. The server still returns meta values as strings
    • CLI --request-sequence, 10 requests in one session, with the replies as paths relative to the JSON file: float32 audio equal bit for bit
    • Server audio is PCM16, so it was compared with the C API's float samples after the same conversion the server applies
  • The same comparison on CUDA, EN F16 and Q4_0, 10 turns each, offline and streamed: both runtimes' /v1/tasks/run and /v1/tasks/stream and the CLI sequence returned the C API's bytes for every turn, and so did 3 JP F16 turns through the parallel runtime
  • Status, 15 cases per runtime, all as expected:
    • 400 invalid_request_error: an unknown kind, bad Base64 and a missing file (each on run and stream), an lfm2_audio.reply with no question before it (three ways), and an artifact sent to VoxCPM1 (loaded; the same request without the artifact returns audio)
    • 200: S2S without artifacts (run and stream), a two-turn control, and an lfm2_audio ASR request with a custom artifact it ignores
    • The CLI exits 1 for the same malformed artifacts and for the reply out of turn, and 0 for the control
  • Of the six refusals, VoxCPM1's ran through the server and YuE2's runs in yue2_request_test. AuK, HeartMuLa, VibeVoice and VoxCPM2 are built, not run
  • The lfm2_audio page's Conversations example, with its placeholders filled in from a real turn 1, returns 200 on both runtimes. With a seed added, it returns the C API's turn 2 byte for byte
  • An EN request for turn 10 is 1.43 MB with the history inline, 95.8 KB with the questions as path and the replies inline, and 4.3 KB with path for both
  • Not run at this head: two conversations at once, and more than one slot (lfm2_audio has one on the CPU)

Earlier revisions, carried over and not rerun at this head. These ran before the rebase, with this branch merged with a pre-merge revision of #838, on an M3 Ultra (Metal) and an A10 (CUDA):

  • Both server runtimes repeated the C API's turns byte for byte: 33 of 33 reply artifacts per runtime on the A10 (EN and JP, offline and streamed, Base64 and path), and 19 of 19 on the M3 (Base64)
  • Both runtimes (EN F16) gave each of 34 requests the expected status. That included 500, as before, for an unknown option and a negative temperature
  • Changing the reader's or YuE2's error type failed a test on both machines. Swapping or removing a kind table entry fails the build
  • Two conversations at once on CUDA, under host load 45 to 90, differed from the C API's replies at a near-tie code in 13 of 26 runs. The cause is a block_reduce race in ggml-cuda, which main has too, and 🚨 fix(ggml-cuda): backport the block_reduce shared-memory race fix (llama.cpp #26385) #837 fixes it. With the fix, 66 concurrent turns per runtime matched, but that was on a quiet host, where a build without the fix also matched
  • parallel_http_live_body_test twice failed to reach its test server, then passed 46 reruns. Its sources outside the test, http.cpp and parallel_http.cpp, are untouched. It passed all 4 runs at this head

Build and run

cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release -DENGINE_BUILD_TESTS=ON -DENGINE_BUILD_EXTENDED_TESTS=ON \
  -DENGINE_BUILD_MODEL_TESTS=ON -DAUDIOCPP_BUILD_C_API=ON -DAUDIOCPP_MODEL_SET=custom \
  -DAUDIOCPP_MODELS=lfm2_audio,auk,heartmula,vibevoice,voxcpm1,voxcpm2,yue2
ninja -C build audiocpp_server audiocpp_cli base64_test cli_request_options_test server_parallel_lifecycle_test \
  server_runtime_selection_test yue2_request_test
(cd build && ctest -R '^(base64_test|cli_request_options_test|server_parallel_lifecycle_test|server_runtime_selection_test|yue2_request_test)$')
curl http://127.0.0.1:8080/v1/tasks/run -H 'Content-Type: application/json' -d '{"model": "my-model",
  "request": {"audio": "/path/to/input.wav", "artifacts": [{"id": "x.tokens", "kind": "acoustic_tokens", "path": "/path/to/tokens.bin"}]}}'

For an lfm2_audio conversation, see Conversations in docs/community_models/lfm2_audio.md.

The CLI request reader is about to decode base64 payloads too, so the
codec moves next to build_info as minitts::app::base64_encode and
base64_decode. The server keeps calling it unqualified through using
declarations. No behaviour change. Its unit test follows it as
base64_test.
A result's artifacts already come back in an "artifacts" array, but no
request route read one, so a client could not send an artifact back.
build_request_from_json now reads an "artifacts" array of objects in the
shape the server writes, {id, kind, payload, meta}, into
TaskRequest::input_artifacts in order. "path" may replace the base64
"payload": the bytes are then read from that file, resolved like the
other request paths. Number and boolean meta values become text as
option values do.

Every user of build_request_from_json gets it: /v1/tasks/run,
/v1/tasks/stream and /v1/tasks/batch on both server runtimes, the CLI
--request-sequence JSON, workflow requests and model_perf. A request
that carried the key used to run with it ignored; now its artifacts
reach the model, and a family that refuses input artifacts turns the
request away.

A malformed entry throws engine::runtime::InvalidRequestError naming
the entry, which both server runtimes answer with HTTP 400 and the CLI
reports as before. The server's max_request_body_bytes bounds inline
payloads, as it does audio_base64. A path must be a regular file, and
the payloads of one request, each /v1/tasks/batch entry being one,
total at most 2 GiB, the default body limit; a static_assert in the
server config keeps the two equal. A batch reads all its entries before
it runs, so with path artifacts it can hold up to 2 GiB per entry.

A kind is read by the name the server writes it with. A static_assert
checks that the reader's table names every ArtifactKind, in enum order,
so a kind added to the enum before Custom, which stays last, fails to
build until the reader names it too. An inline payload decodes straight
into the artifact's bytes through the new base64_decode_bytes, a
std::byte twin of base64_decode, so the decoded bytes are not held
twice.

The C API comment that cited a line of request.cpp now names
build_request_from_cli instead, as the line moved.

Tests: the CLI request test reads inline, file and data URI payloads,
every kind name the server writes and the rejections. base64_test
checks base64_decode_bytes against base64_decode. The parallel
lifecycle fixture now returns its input artifacts as its result's, so
run, stream and batch requests are checked end to end, along with their
400s. The legacy runtime test checks the 400s and that valid artifacts
get as far as the model load.
AuK, HeartMuLa, VibeVoice, VoxCPM1, VoxCPM2 and YuE2 refuse any input
artifact with a std::runtime_error. Since the previous commit a task
request's "artifacts" reach the model, so a request to one of them that
carries an artifact, which used to run with it ignored, now fails, and
through the server as a 500 server_error. The request is what is wrong,
so these six checks throw InvalidRequestError instead, which both
server runtimes answer with 400. The CLI and the C API report it as
before, since it is still a std::runtime_error.

Only the artifact checks change. The same functions' other refusals,
such as audio_input to HeartMuLa or YuE2, could be reached before and
stay as they are.

yue2_request_test checks YuE2's refusal, the one of the six a test can
reach without weights. The server README names the six.
The lfm2_audio page said the server takes no request artifacts, so that
through it each S2S request was a first turn and conversations ran
through the C API only. With request artifacts the server and the CLI
take the history too. The Conversations section now shows a later turn
through /v1/tasks/run, with the earlier question as a path and the
reply's payload and meta as the result returned them, says the CLI's
request JSON takes the same array, and that the CLI's payload_hex
results go back as a file's path. The live route still takes no
artifacts, which the limitation now says instead. The server README
now says that a relative artifact path resolves against the server's
working directory.
@0xShug0
0xShug0 merged commit 71268e5 into 0xShug0:main Oct 9, 2026
11 checks passed
@0xShug0

0xShug0 commented Oct 9, 2026

Copy link
Copy Markdown
Owner

Thanks @ykhrustalev This is a great transport improvement 🎉 And thank you for clearly acknowledging its limitations.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants