Repository navigation
feat(lfm2_audio): multi-turn speech-to-speech with client-carried history - #838
Merged
Merged
Conversation
A conversation's next turn replays the earlier replies. liquid-audio's _prefill puts their audio frames into the prompt as the sum of each frame's codebook embeddings, the input its decode step feeds, but Lfm2Prompt could hold only text ids and audio-in positions. Add frame_positions and frame_codes, and embed those frames in the prefill. The sum over codebooks is now one helper that the decode graph and the prefill both build, so a frame replayed in a prompt sums exactly as step_audio summed it. For one frame it builds the decode graph's ops as before, plus a reshape that copies nothing. A prompt without frames builds the prefill graph it built before, so ASR, TTS and S2S output does not change. start() rejects frames on a backbone loaded without the audio embedding, frames short of a code per codebook, codes outside the codebook, and frame positions that do not increase, leave the prompt or take an audio position. The backbone test fixture gains a small audio embedding (2 codebooks of 5 codes). A reply of text tokens and frames prefilled after a question gives the logits of feeding it step by step and of the double precision reference.
A conversation's next turn prefills its whole history, thousands of steps after a few turns, and the one-shot prefill holds prompt^2 attention scores per head: 2 GiB at 4096 steps with the 32 heads of the checkpoints. start() gains an Lfm2Prefill argument after the cache policy. OneShot, the default, is the prefill as before, so ASR, TTS and S2S do not change. Chunked runs the prompt in blocks of 256 positions at multiples of 256: each block writes its keys and values straight into the decode cache, and its short-conv layers carry the conv state over in the decode graph's conv tails, as decode steps do. A block's attention layers run build_with_static_cache_block on a view of the cache's first rows, up to the block's own end, rather than on the whole cache. What a block computes then depends on the prompt up to its end and not on the cache's length, which follows max_steps. Over the whole cache it would not: ggml's CPU flash attention splits one query's keys among the threads from 512 keys on, so a block of one position would change in its last bits with max_steps. Every block has its own graph, and they share one allocator kept with the backbone. Chunked and one-shot logits agree to float rounding, not bit for bit: blocks use the flash attention that decode steps use. On a small synthetic model with 64-wide heads they agree to within 5e-7 of the largest logit on the CPU, and to within 6e-4 on Metal and CUDA, whose flash attention rounds part of its input to half precision. On the EN F16 checkpoint, from 256 to 8192 steps, they agree to within 1.1% of the largest logit on CUDA, 0.15% on Metal, 0.07% on an x86 CPU and 0.6% on an M3 CPU, with the same argmax. The tests prefill a 600-step prompt of text, audio rows across the first block edge and a reply of tokens and frames across the second. Cut on either side of each edge, it matches the double precision reference and the one-shot prefill when chunked, and so do the steps after it. A 340-step reply prefilled chunked matches feeding it step by step. Prefills into caches of 768 and 1536 steps give the same bits, and so does a chunked request after longer, shorter, one-shot and text requests and in a cache another prompt filled. Both prefills reject bad frames, NaN audio and requests past the context.
liquid-audio's ChatState keeps every turn of a conversation, and each reply prefills all of it again. make_lfm2_chat_prompt builds that sequence for a turn after earlier ones, as the README's multi-turn example holds it: the system prompt once, each earlier question's audio positions, each reply step by step with <|im_end|>\n<|im_start|>user\n after it, then the new question. demo/chat.py, as written, adds the system turn again on every turn, as its len(chat.text) == 1 check always holds; this prompt keeps it once. A reply goes in as the token ids and frames the generator yielded, the end-of-audio frame included and a reply cut off at max_tokens kept as it is, as ChatState keeps them, so no text is tokenized again. Without history it is the prompt S2S already builds. A request is to carry its earlier turns in its input artifacts: for each turn in order, an lfm2_audio.question with the question's WAV bytes (kind Custom) and the lfm2_audio.reply artifact that turn's result returned. The reply artifact is kind AcousticTokens, with a payload of little-endian int32 values: a token id per text step, and -1 followed by one code per codebook per frame, about 470 bytes per second of reply audio. Its meta names the format, lfm2_audio.reply/1, and the checkpoint's language, codebook count, codebook size and vocabulary size, which must match where it is read, and holds the step count and whether the reply ended. It carries ids and codes only, so it replays on any quantization of the checkpoint. read_lfm2_conversation reads that history. It rejects artifacts it does not know, artifacts out of turn, a question without its reply, a WAV that does not read, and a reply whose kind, format, meta, length, tokens or codes are not what make_lfm2_reply_artifact writes for the checkpoint. A conversation's prompt and max_tokens may take 8192 steps together, and require_lfm2_conversation_room throws a CapacityError past that, saying to leave out the oldest turns. read_lfm2_conversation counts each reply from its meta before it reads the payload, and stops at the first reply that leaves no room for max_tokens, so a history over the limit is not read in full only to be turned away. lfm2_chat_prompt_text_steps counts the prompt's steps other than the questions' audio, so that the limit can be checked before any question is encoded. Nothing calls these yet. lfm2_audio_chat_test checks the prompt against ChatState's sequence for zero, one and two earlier turns on the byte vocabulary, the text steps of each, the artifact's bytes and its round trip, the limit, and each rejection.
An S2S request answered one question as a new conversation. Its input artifacts now carry the earlier turns, an lfm2_audio.question and an lfm2_audio.reply for each, and with the new request option return_codes the result, offline or streamed, also returns the reply as an lfm2_audio.reply artifact to send back with the next turn. The session keeps nothing between requests. Each turn prefills the whole conversation again, as liquid-audio's chat demo does: each earlier question goes through the encoder on its own, the embeddings go into the chat prompt in order, and the earlier replies go in as their tokens and frames. A turn with history takes the chunked prefill. A turn without it builds the prompt it always did and takes the one-shot prefill, so a request without artifacts or return_codes gives the bytes it gave before. The reply artifact holds every step the generator yielded, the end-of-audio frame included. A reply cut off at max_tokens, and a stream finished before its reply is over, return the steps so far with ended=false. A conversation's prompt and max_tokens may take 8192 steps together; past that a turn fails with a CapacityError that says to leave out the oldest turns, and no turn is ever dropped silently. The limit is checked as soon as the request shows it passed: when the request is read, on its replies and chat markup, so that a stream fails before its audio comes, then after each earlier question is encoded, and last on the whole prompt. A first turn is held only to the context, as before. Earlier questions are held to lfm2_audio.max_pass_seconds like the new one when the request is read, so a stream with a bad history fails before its audio comes. ASR and TTS turn away return_codes and lfm2_audio.* artifacts, and S2S turns away any other artifact. The model spec gains return_codes.
The S2S session had no unit test that runs a reply, as the synthetic packages had no vocoder or detokenizer. lfm2_audio_test_package.h gains writers for both, and lfm2_audio_chat_test builds an S2S package on them. Its backbone is random but leaves the dimensions its inputs need to those inputs, so every reply writes "hi", speaks seven frames whose first codes follow a fixed path, and ends on the end-of-audio frame. One head of the last attention layer averages a mark that only the end-of-audio frames of earlier replies carry into the logits of the second codebook, so earlier turns change its codes. On that package, offline: a first turn with and without return_codes, the reply artifact's steps and meta, a second and a third turn with the history the turns before returned, which changes the reply, each turn again, in another order and in a fresh session, to the bit, a reply cut at max_tokens and a turn after it, and the rejections: artifacts of another id or out of turn, a reply of the other checkpoint, return_codes that is not a boolean, a conversation over 8192 steps, turned away by each check from its replies to its whole prompt, next to a turn exactly at the limit, while a first turn is not held to it, an earlier question over lfm2_audio.max_pass_seconds, and history or return_codes on ASR and TTS. Streamed: each turn returns the offline turn's reply artifact, again to the bit, a stream finished early or before its reply returns the steps so far as not ended, and a bad history, or one whose text passes the limit, fails at start_stream. test_lfm2_audio_s2s gains a conversation on the published package: the first reply's artifact holds its text and speech, a second turn (assets/resources/a.wav) sent with the first gets another reply than without it and the same bytes when sent again, and a third turn replays both.
The docs said S2S answers one turn per request, as a new conversation. A Conversations section now says how a request carries the turns before it: the lfm2_audio.question and lfm2_audio.reply artifacts in conversation order, the reply artifact's payload and meta, that a turn depends on its request alone, the 8192-step limit and when a request over it fails, how to send the history through the C API, and that the server returns the reply artifact but does not take request artifacts yet. The prompt is the one ChatState holds in liquid-audio's README multi-turn example, with the system prompt once. The section says that demo/chat.py, as written, adds the system turn again on every turn and that audio.cpp does not, as in ten-turn English conversations run with liquid-audio the repeated turn left replies answering 1 of 21 questions about earlier turns, against 7 of 21. It also gives the range checked, 10 turns with prompts of up to about 5000 steps, in which replies used the earlier turns but not what users said about themselves and late Japanese replies sometimes repeated an earlier one, and says to leave out the oldest turns past it. The request options list return_codes, the single-turn bullet leaves the limitations, and the validation section covers the new tests.
The model spec gained the return_codes request option, so the page is rebuilt with npm ci and npm run build in webui/native from this branch's webui/ and model_specs/. It differs from the base's page only in that option's entry and the bundle's hashed name. Built the same way, the base's sources (cf124a6) give its committed page byte for byte.
Of the errors a session throws, only CapacityError reaches a client as a 400. Any other, a std::runtime_error or a std::invalid_argument alike, gets 500 server_error from both server runtimes, also when the request is at fault: an input the model does not take, or one that does not read. The caller then looks for a server fault that is not there. std::invalid_argument would not tell them apart, as sessions throw it for their own faults too. engine::runtime::InvalidRequestError, next to CapacityError, is for those. It is a std::runtime_error, so the CLI and the C API report it as before, and both server runtimes answer it with 400 invalid_request_error. CapacityError stays as it is: there the request is fine, only too big for the device. Nothing throws the new error yet. The parallel lifecycle test gets a session that turns its request away and checks the 400 on /v1/tasks/run, offline and streaming, and on /v1/tasks/stream, with the lease released each time, and that another failure still leaves the handler for the transport's 500.
Through the server, history that S2S turns away (out of turn, another format, checkpoint or kind, bad meta or payload, a question that does not read as WAV), a return_codes that is neither true nor false, and history or return_codes sent to ASR or TTS came back as 500 server_error. They are the caller's to fix, so they now throw InvalidRequestError, which both server runtimes answer with 400. So does reject_options, which turns return_codes away on ASR and TTS, and with it the other options a task does not take, a 500 on main too. A reply step about to be written that does not fit the checkpoint stays a std::runtime_error, a fault of the session, so check_step takes the error type. The conversation limit stays a CapacityError. lfm2_audio's other checks, and the option parsers every family shares, are unchanged. lfm2_audio_chat_test now requires InvalidRequestError for each request it turns away, and an error of another type where a prompt or a reply artifact is built from bad steps. The docs say which errors are a 400.
Owner
|
OOTO now will check later today... |
Owner
|
@ykhrustalev PR merged! One thing can be addressed in a follow-up PR: use #834’s packed-buffer-aware validator in the new BlockGraph, consistent with the existing prefill and decode graphs. But no actual failure from this omission was found. |
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Multi-turn S2S for
lfm2_audio, with the conversation carried by the client.Problem
ChatStateprefills every earlier turn for each replyWhat this PR changes
make_lfm2_chat_promptfollows liquid-audio's README multi-turn example: the system prompt once, each earlier question and its reply as generated (cut replies included), then the new question.demo/chat.pyrepeats the system turn every turn (len(chat.text) == 1always holds); this PR does not, on purposeTaskRequest.input_artifacts: per earlier turn anlfm2_audio.question(WAV bytes) and thelfm2_audio.replyit returned; the session keeps nothing. Optionreturn_codesreturns that reply, offline or streamed (kindacoustic_tokens: int32 token ids and frames), so it replays on any quantizationended=truewithout a final end-of-audio frame, is taken as sent. ASR and TTS rejectreturn_codesandlfm2_audio.*artifactsengine::runtime::InvalidRequestError(errors.h) for a request the caller can fix; both server runtimes answer it with 400invalid_request_error. lfm2_audio throws it for the rejections above, a non-booleanreturn_codesand options a task does not take. The request-artifacts PR carries the same commit, so either can merge firstmax_tokensmay take 8192 steps; past that the turn fails withCapacityErroras soon as the request shows it. Nothing is dropped silently. A first turn keeps the context's limit, as beforeWhat changes for users
return_codesgive main's bytes. S2S now turns away input artifacts other thanlfm2_audio.questionandlfm2_audio.reply, which it ignored before; ASR and TTS turn awaylfm2_audio.*onespayload_hexartifact JSON as is; without it the server ignores history in a request and answers a first turn. The second to land updates the docs' server notes. The live route stays single-turnreturn_codesor an option a task does not take, such astemperatureon ASR (a 500 on main). A negative temperature and an unknown option stay 500sTwo conversations at once on CUDA
Two conversations at once in one CUDA process sometimes gave other bytes than each alone, from a near-tie code on: 13 of 26 server runs under host load 45 to 90, and 1 of 162 turns of two C API sessions at this head's code on a quiet host. The diverged turn we traced first differed at the encoder's
soft_max_f32, which hits a shared-memory race in ggml-cuda'sblock_reducethat main has too; #837 (a backport of llama.cpp #26385) fixes it, and with it 0 of 159 such C API turns differed. The server runs under load were not repeated with the fix. Without the fix, single requests repeated byte for byte in every normal run (below), but racecheck finds the race with one session too, and under memcheck's timing a single-turn reply on main differed in 5 of 5 repeats. Metal, which runs none of this code, never diverged. Once, a traced two-session run without the fix stopped on an illegal memory access in a kernel the error did not name; it did not recur in any later run.Testing
At this head (base cf124a6; merges cleanly with main 75d0294 and the request-artifacts PR), on an M3 Ultra (Metal) and an A10 (CUDA):
parallel_http_live_body_test, whose own sources neither PR touches, failed twice to reach its test server in earlier runs and passed 46 rerunstest_lfm2_audio_s2s(EN F16, 3 turns) passes on CUDA and Metal with the same later-turn textspathartifacts on the streamed one), JP offline (10) and streamed (3), 33 of 33 reply artifacts per runtime. M3: EN offline beside JP streamed (13 turns), then EN streamed and JP offline in turn (6), 19 of 19 per runtime in artifact, meta, text and audioAt an earlier revision lacking only the request error, against liquid-audio's own 3-turn conversations (JP and EN, two layouts, up to 2272 steps): the prompt equals the dumped
ChatStateposition by position, but for the demo's repeated system turns. Teacher-forced, JP F32 on the CPU differs in 0 of 93656 picks on the M3 and 1 on x86 (a 3e-6 tie), each turn counted once per prefill; the three greedy README-layout turns through the session match exactly on both CPUs. EN F16 on CUDA agrees with fp32 on 99.96% of 106568 picks, counted the same way; the other 42 are near-ties (within 0.0096 in our logits, 0.026 in the reference's), 20 chunked against 22 one-shot, 12 against 10 on the later turns, where the session prefills chunked.For the request error, at an earlier revision with this head's code: tests on both machines catch a missing server catch or a swapped error type, and 44 requests through each server runtime on both machines (Metal, CUDA) got the expected status, still 500 for an unknown option and a negative temperature.
Before the rebase onto cf124a6, same prefill, prompt and artifacts: single-turn output matched main's bytes in 27 cases on the A10 (CUDA, CPU) and 32 on the M3. C API conversations of 3 and 10 turns (EN F16, Q8_0, Q4_0 and JP F16 on CUDA and Metal, JP F32 on the CPU) repeated and replayed bitwise, streamed turns returned the offline artifact, records replayed across quantizations and devices, and 149 server and 39 CLI turns through the request-artifacts PR, one conversation at a time, matched them. Chunked and one-shot logits (EN F16, 256 to 8192 steps) agree within 1.1% of the largest logit on CUDA, 0.15% on Metal, 0.07% on the x86 CPU and 0.6% on the M3 CPU, with the same argmax.
Quality, with liquid-audio itself (540 replies: 10-turn EN and JP conversations in two layouts, plus each question alone): every reply kept its audio, and the median WER of reply text against its own audio did not grow with turns. With the system prompt once, JP turn 8 repeated turn 3 in 3 of 9 conversations and two later replies hit max_new_tokens; EN answered 7 of 21 questions about earlier turns (1 with the demo's layout); questions about the user's own details failed either way.
Prefill, EN F16, one-shot / chunked, warm time and peak over the loaded model, before the rebase with this head's prefill code and ggml. A10 at host load 53 to 69. M3: the middle of three runs at load 6 to 20, within 9% of each other; a fourth, at load 3 and 20 to 60% slower, is left out:
Reply artifacts take about 470 B per second of audio; a 10-turn history is about 1 MB, mostly questions.
Build and run