Repository navigation
feat: input artifacts in server and CLI task requests - #844
Merged
Merged
Conversation
The CLI request reader is about to decode base64 payloads too, so the codec moves next to build_info as minitts::app::base64_encode and base64_decode. The server keeps calling it unqualified through using declarations. No behaviour change. Its unit test follows it as base64_test.
A result's artifacts already come back in an "artifacts" array, but no
request route read one, so a client could not send an artifact back.
build_request_from_json now reads an "artifacts" array of objects in the
shape the server writes, {id, kind, payload, meta}, into
TaskRequest::input_artifacts in order. "path" may replace the base64
"payload": the bytes are then read from that file, resolved like the
other request paths. Number and boolean meta values become text as
option values do.
Every user of build_request_from_json gets it: /v1/tasks/run,
/v1/tasks/stream and /v1/tasks/batch on both server runtimes, the CLI
--request-sequence JSON, workflow requests and model_perf. A request
that carried the key used to run with it ignored; now its artifacts
reach the model, and a family that refuses input artifacts turns the
request away.
A malformed entry throws engine::runtime::InvalidRequestError naming
the entry, which both server runtimes answer with HTTP 400 and the CLI
reports as before. The server's max_request_body_bytes bounds inline
payloads, as it does audio_base64. A path must be a regular file, and
the payloads of one request, each /v1/tasks/batch entry being one,
total at most 2 GiB, the default body limit; a static_assert in the
server config keeps the two equal. A batch reads all its entries before
it runs, so with path artifacts it can hold up to 2 GiB per entry.
A kind is read by the name the server writes it with. A static_assert
checks that the reader's table names every ArtifactKind, in enum order,
so a kind added to the enum before Custom, which stays last, fails to
build until the reader names it too. An inline payload decodes straight
into the artifact's bytes through the new base64_decode_bytes, a
std::byte twin of base64_decode, so the decoded bytes are not held
twice.
The C API comment that cited a line of request.cpp now names
build_request_from_cli instead, as the line moved.
Tests: the CLI request test reads inline, file and data URI payloads,
every kind name the server writes and the rejections. base64_test
checks base64_decode_bytes against base64_decode. The parallel
lifecycle fixture now returns its input artifacts as its result's, so
run, stream and batch requests are checked end to end, along with their
400s. The legacy runtime test checks the 400s and that valid artifacts
get as far as the model load.
AuK, HeartMuLa, VibeVoice, VoxCPM1, VoxCPM2 and YuE2 refuse any input artifact with a std::runtime_error. Since the previous commit a task request's "artifacts" reach the model, so a request to one of them that carries an artifact, which used to run with it ignored, now fails, and through the server as a 500 server_error. The request is what is wrong, so these six checks throw InvalidRequestError instead, which both server runtimes answer with 400. The CLI and the C API report it as before, since it is still a std::runtime_error. Only the artifact checks change. The same functions' other refusals, such as audio_input to HeartMuLa or YuE2, could be reached before and stay as they are. yue2_request_test checks YuE2's refusal, the one of the six a test can reach without weights. The server README names the six.
The lfm2_audio page said the server takes no request artifacts, so that through it each S2S request was a first turn and conversations ran through the C API only. With request artifacts the server and the CLI take the history too. The Conversations section now shows a later turn through /v1/tasks/run, with the earlier question as a path and the reply's payload and meta as the result returned them, says the CLI's request JSON takes the same array, and that the CLI's payload_hex results go back as a file's path. The live route still takes no artifacts, which the limitation now says instead. The server README now says that a relative artifact path resolves against the server's working directory.
Owner
|
Thanks @ykhrustalev This is a great transport improvement 🎉 And thank you for clearly acknowledging its limitations. |
This was referenced Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Server and CLI task requests can carry input artifacts, in the shape results already return them in. With lfm2_audio's multi-turn S2S (#838) on main, this lets server and CLI clients hold a conversation.
Problem
artifactsarray, but no request route reads one. Only the C API fillsTaskRequest::input_artifacts, so server and CLI clients cannot send back what a result gave themlfm2_audio.questionandlfm2_audio.reply, and withreturn_codesit returns each turn'slfm2_audio.reply. The server returns that reply but cannot take it back. So through the server and the CLI, every S2S request is a first turn, and the lfm2_audio docs point conversations to the C APIWhat this PR changes
build_request_from_jsonreads anartifactsarray, in order, intoTaskRequest::input_artifacts. That covers/v1/tasks/run,/v1/tasks/streamand each/v1/tasks/batchentry on both server runtimes, plus the CLI's--request-sequenceJSON and workflow requestsid(ids may repeat), akindas the server names it, a Base64payloador data URI, and optionalmeta. Numbers and booleans inmetabecome text, asoptionsvalues do (1.0becomes1).pathmay replacepayload, and it resolves the way the request'saudiodoes: on the server against its working directory, in the CLI against the JSON fileengine::runtime::InvalidRequestError, which main already has, with a message naming the entry, such asartifacts[1] (example.tokens): unknown kind 'tokens'. Both server runtimes answer it with 400invalid_request_error, and the CLI exits 1max_request_body_bytesbounds inline payloads, as it boundsaudio_base64. One request's payloads, inline and from files, may total 2 GiB, the default body limit. Astatic_assertkeeps the two equalstatic_assertchecks that the kind table names everyArtifactKindin enum order, up toCustom, which stays last. Inline payloads decode straight into the artifact (newbase64_decode_bytes)minitts::app) so the CLI can use it.server_base64_testbecomesbase64_test/v1/tasks/run, with the earlier question as apathand the reply's payload and meta as the result returned them. It also says the CLI's request JSON takes the same array, and that the CLI'spayload_hexresults go back as a file'spath. The limitation now names only the live route. The server README now says what a relativepathresolves againstWhat changes for users, and what does not
artifacts, and which ones a model reads is up to its family. lfm2_audio S2S reads a conversation's history, so server and CLI clients can now carry one: for each earlier turn, in order, the question's WAV bytes aslfm2_audio.question(kindcustom), then thelfm2_audio.replythat turn returned. The rules are feat(lfm2_audio): multi-turn speech-to-speech with client-carried history #838's: S2S turns away other ids and a history out of order, while ASR and TTS turn awaylfm2_audio.*ids and ignore the restartifactsare unchanged/v1/audio/speech/live, takes no artifacts, so each of its requests is still a first turnartifactskey used to be ignored and now gets a 400. So does a non-empty one sent to the six families, which used to run (the CLI exits 1 for both). So does a history that lfm2_audio turns away; before, the server dropped the key and answered a first turnserver_error, as before: a body overmax_request_body_bytes, and request errors that throw a plainstd::runtime_error, such as an unknown option or an option value the shared parsers reject. In this PR, only the request reader and the six refusals throwInvalidRequestErrorpathreads any regular file the server process can read, asaudioandvoice_refdo. No model echoes its input artifacts. If anlfm2_audio.reply'spathnames some other file, the refusal message can quote an out-of-range 32-bit value from that file, or its sizemax_request_body_bytes. Each batch entry is one request, and a batch reads all its entries before it runs, so withpathartifacts a batch can hold 2 GiB per entrypayload_hex, which requests do not take. Write the bytes to a file, or Base64 themTesting
At this head, on an A10 host: its CPU (a Xeon 8358, 8 threads) and its A10 (CUDA). Release builds with gcc 12, with lfm2_audio and the six families built in. Main still has a ggml-cuda race that can change a reply at a near-tie code (#837 fixes it), so the CUDA runs used a build with #837's fix merged in. Metal ran only at an earlier revision (below).
lfm2_audio_encoder_cpu_repack_testincluded. The skipped two are the C API model tests, which need a model root. The two sets differ only by the newyue2_request_testand thebase64_testrename. Three more runs of the same filtered set on each tree passed,parallel_http_live_body_testincluded. The PR's unit tests cover every payload form and kind name, the rejections other than file I/O failures, the 2 GiB file total (skipped on Windows), run, stream and batch end to end on the parallel runtime, and the legacy runtime's 400sreturn_codesand a fixed seed per turn. Each server and CLI turn sent the C API's earlier questions and replies as history. Each one returned the C API's reply artifact (payload, kind and meta), text and audio byte for byte, 384 of 384 checks:/v1/tasks/run(questions aspath, replies Base64) and/v1/tasks/stream(both Base64, also checking the events' audio and partial text), 10 turns each--parallel-jobs, one slot): EN/v1/tasks/run(both Base64) and/v1/tasks/stream(bothpath), 10 turns each, and JP/v1/tasks/run, 3 turnsmetasent back as JSON numbers and booleans, 3 turns per runtime: same replies. The server still returnsmetavalues as strings--request-sequence, 10 requests in one session, with the replies as paths relative to the JSON file: float32 audio equal bit for bit/v1/tasks/runand/v1/tasks/streamand the CLI sequence returned the C API's bytes for every turn, and so did 3 JP F16 turns through the parallel runtimeinvalid_request_error: an unknown kind, bad Base64 and a missing file (each on run and stream), anlfm2_audio.replywith no question before it (three ways), and an artifact sent to VoxCPM1 (loaded; the same request without the artifact returns audio)customartifact it ignoresyue2_request_test. AuK, HeartMuLa, VibeVoice and VoxCPM2 are built, not runseedadded, it returns the C API's turn 2 byte for bytepathand the replies inline, and 4.3 KB withpathfor bothEarlier revisions, carried over and not rerun at this head. These ran before the rebase, with this branch merged with a pre-merge revision of #838, on an M3 Ultra (Metal) and an A10 (CUDA):
path), and 19 of 19 on the M3 (Base64)block_reducerace in ggml-cuda, which main has too, and 🚨 fix(ggml-cuda): backport the block_reduce shared-memory race fix (llama.cpp #26385) #837 fixes it. With the fix, 66 concurrent turns per runtime matched, but that was on a quiet host, where a build without the fix also matchedparallel_http_live_body_testtwice failed to reach its test server, then passed 46 reruns. Its sources outside the test, http.cpp and parallel_http.cpp, are untouched. It passed all 4 runs at this headBuild and run
For an lfm2_audio conversation, see Conversations in docs/community_models/lfm2_audio.md.