Skip to content

feat(lfm2_audio): multi-turn speech-to-speech with client-carried history - #838

Merged
0xShug0 merged 10 commits into
0xShug0:mainfrom
Liquid4All:lfm2-audio-multiturn
Oct 8, 2026
Merged

0xShug0 merged 10 commits into
0xShug0:mainfrom
Liquid4All:lfm2-audio-multiturn

Conversation

@ykhrustalev

Copy link
Copy Markdown
Contributor

Multi-turn S2S for lfm2_audio, with the conversation carried by the client.

Problem

  • Every S2S request started a new conversation, so a follow-up could not refer to an earlier answer. liquid-audio's ChatState prefills every earlier turn for each reply
  • The one-shot prefill holds prompt^2 attention scores per head: 2 GiB at 4096 steps across the 32 heads

What this PR changes

  • Backbone: a prefill takes earlier reply frames, embedded as in decoding. A chunked prefill writes 256-step blocks straight into the decode cache, each attending only up to its own end, so the cache size does not change its output. ASR, TTS and first turns keep the one-shot prefill
  • make_lfm2_chat_prompt follows liquid-audio's README multi-turn example: the system prompt once, each earlier question and its reply as generated (cut replies included), then the new question. demo/chat.py repeats the system turn every turn (len(chat.text) == 1 always holds); this PR does not, on purpose
  • History comes in TaskRequest.input_artifacts: per earlier turn an lfm2_audio.question (WAV bytes) and the lfm2_audio.reply it returned; the session keeps nothing. Option return_codes returns that reply, offline or streamed (kind acoustic_tokens: int32 token ids and frames), so it replays on any quantization
  • Malformed, foreign or out-of-turn artifacts are rejected by name. A reply is checked step by step against the checkpoint, not as a whole: an empty reply, or ended=true without a final end-of-audio frame, is taken as sent. ASR and TTS reject return_codes and lfm2_audio.* artifacts
  • New engine::runtime::InvalidRequestError (errors.h) for a request the caller can fix; both server runtimes answer it with 400 invalid_request_error. lfm2_audio throws it for the rejections above, a non-boolean return_codes and options a task does not take. The request-artifacts PR carries the same commit, so either can merge first
  • With history, prompt plus max_tokens may take 8192 steps; past that the turn fails with CapacityError as soon as the request shows it. Nothing is dropped silently. A first turn keeps the context's limit, as before
  • Tests: the synthetic package gains a vocoder and detokenizer, so a unit test runs an S2S reply. Docs: a Conversations section. The seventh commit only rebuilds the WebUI page

What changes for users

  • Requests without artifacts or return_codes give main's bytes. S2S now turns away input artifacts other than lfm2_audio.question and lfm2_audio.reply, which it ignored before; ASR and TTS turn away lfm2_audio.* ones
  • Conversations work through the C API. The server returns reply artifacts but takes history only with the request-artifacts PR (feat: input artifacts in server and CLI task requests), which does not take the CLI's payload_hex artifact JSON as is; without it the server ignores history in a request and answers a first turn. The second to land updates the docs' server notes. The live route stays single-turn
  • Through the server an lfm2_audio rejection is a 400 with its message: with this PR alone a bad return_codes or an option a task does not take, such as temperature on ASR (a 500 on main). A negative temperature and an unknown option stay 500s
  • The docs say up to 10 turns were checked; drop the oldest past that

Two conversations at once on CUDA
Two conversations at once in one CUDA process sometimes gave other bytes than each alone, from a near-tie code on: 13 of 26 server runs under host load 45 to 90, and 1 of 162 turns of two C API sessions at this head's code on a quiet host. The diverged turn we traced first differed at the encoder's soft_max_f32, which hits a shared-memory race in ggml-cuda's block_reduce that main has too; #837 (a backport of llama.cpp #26385) fixes it, and with it 0 of 159 such C API turns differed. The server runs under load were not repeated with the fix. Without the fix, single requests repeated byte for byte in every normal run (below), but racecheck finds the race with one session too, and under memcheck's timing a single-turn reply on main differed in 5 of 5 repeats. Metal, which runs none of this code, never diverged. Once, a traced two-session run without the fix stopped on an illegal memory access in a kernel the error did not name; it did not recur in any later run.

Testing
At this head (base cf124a6; merges cleanly with main 75d0294 and the request-artifacts PR), on an M3 Ultra (Metal) and an A10 (CUDA):

  • ctest, on this branch and merged with the request-artifacts PR: 28 of 28 on the M3, 29 of 29 on the A10. parallel_http_live_body_test, whose own sources neither PR touches, failed twice to reach its test server in earlier runs and passed 46 reruns
  • test_lfm2_audio_s2s (EN F16, 3 turns) passes on CUDA and Metal with the same later-turn texts
  • Merged with the request-artifacts PR, both server runtimes repeat earlier C API conversations byte for byte. A10, one at a time: EN offline and streamed (10 turns each, path artifacts on the streamed one), JP offline (10) and streamed (3), 33 of 33 reply artifacts per runtime. M3: EN offline beside JP streamed (13 turns), then EN streamed and JP offline in turn (6), 19 of 19 per runtime in artifact, meta, text and audio

At an earlier revision lacking only the request error, against liquid-audio's own 3-turn conversations (JP and EN, two layouts, up to 2272 steps): the prompt equals the dumped ChatState position by position, but for the demo's repeated system turns. Teacher-forced, JP F32 on the CPU differs in 0 of 93656 picks on the M3 and 1 on x86 (a 3e-6 tie), each turn counted once per prefill; the three greedy README-layout turns through the session match exactly on both CPUs. EN F16 on CUDA agrees with fp32 on 99.96% of 106568 picks, counted the same way; the other 42 are near-ties (within 0.0096 in our logits, 0.026 in the reference's), 20 chunked against 22 one-shot, 12 against 10 on the later turns, where the session prefills chunked.

For the request error, at an earlier revision with this head's code: tests on both machines catch a missing server catch or a swapped error type, and 44 requests through each server runtime on both machines (Metal, CUDA) got the expected status, still 500 for an unknown option and a negative temperature.

Before the rebase onto cf124a6, same prefill, prompt and artifacts: single-turn output matched main's bytes in 27 cases on the A10 (CUDA, CPU) and 32 on the M3. C API conversations of 3 and 10 turns (EN F16, Q8_0, Q4_0 and JP F16 on CUDA and Metal, JP F32 on the CPU) repeated and replayed bitwise, streamed turns returned the offline artifact, records replayed across quantizations and devices, and 149 server and 39 CLI turns through the request-artifacts PR, one conversation at a time, matched them. Chunked and one-shot logits (EN F16, 256 to 8192 steps) agree within 1.1% of the largest logit on CUDA, 0.15% on Metal, 0.07% on the x86 CPU and 0.6% on the M3 CPU, with the same argmax.

Quality, with liquid-audio itself (540 replies: 10-turn EN and JP conversations in two layouts, plus each question alone): every reply kept its audio, and the median WER of reply text against its own audio did not grow with turns. With the system prompt once, JP turn 8 repeated turn 3 in 3 of 9 conversations and two later replies hit max_new_tokens; EN answered 7 of 21 questions about earlier turns (1 with the demo's layout); questions about the user's own details failed either way.

Prefill, EN F16, one-shot / chunked, warm time and peak over the loaded model, before the rebase with this head's prefill code and ggml. A10 at host load 53 to 69. M3: the middle of three runs at load 6 to 20, within 9% of each other; a fourth, at load 3 and 20 to 60% slower, is left out:

Steps A10 ms A10 MiB M3 Metal ms M3 Metal MiB
4096 589 / 280 3018 / 310 620 / 535 2905 / 249
8192 1845 / 636 10052 / 528 1760 / 1140 9885 / 459

Reply artifacts take about 470 B per second of audio; a 10-turn history is about 1 MB, mostly questions.

Build and run

cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release -DENGINE_BUILD_TESTS=ON \
  -DENGINE_BUILD_MODEL_TESTS=ON -DAUDIOCPP_MODEL_SET=custom -DAUDIOCPP_MODELS=lfm2_audio
ninja -C build lfm2_audio_chat_test lfm2_audio_backbone_test test_lfm2_audio_s2s
(cd build && ctest -R 'lfm2_audio_(chat|backbone)_test')
./build/bin/test_lfm2_audio_s2s  # with lfm2_audio_1_5b_f16 in models/

A conversation's next turn replays the earlier replies. liquid-audio's
_prefill puts their audio frames into the prompt as the sum of each
frame's codebook embeddings, the input its decode step feeds, but
Lfm2Prompt could hold only text ids and audio-in positions. Add
frame_positions and frame_codes, and embed those frames in the prefill.

The sum over codebooks is now one helper that the decode graph and the
prefill both build, so a frame replayed in a prompt sums exactly as
step_audio summed it. For one frame it builds the decode graph's ops as
before, plus a reshape that copies nothing. A prompt without frames
builds the prefill graph it built before, so ASR, TTS and S2S output
does not change.

start() rejects frames on a backbone loaded without the audio
embedding, frames short of a code per codebook, codes outside the
codebook, and frame positions that do not increase, leave the prompt or
take an audio position.

The backbone test fixture gains a small audio embedding (2 codebooks of
5 codes). A reply of text tokens and frames prefilled after a question
gives the logits of feeding it step by step and of the double precision
reference.
A conversation's next turn prefills its whole history, thousands of
steps after a few turns, and the one-shot prefill holds prompt^2
attention scores per head: 2 GiB at 4096 steps with the 32 heads of
the checkpoints. start() gains an Lfm2Prefill argument after the cache
policy. OneShot, the default, is the prefill as before, so ASR, TTS
and S2S do not change. Chunked runs the prompt in blocks of
256 positions at multiples of 256: each block writes its keys and
values straight into the decode cache, and its short-conv layers carry
the conv state over in the decode graph's conv tails, as decode steps
do.

A block's attention layers run build_with_static_cache_block on a view
of the cache's first rows, up to the block's own end, rather than on
the whole cache. What a block computes then depends on the prompt up to
its end and not on the cache's length, which follows max_steps. Over
the whole cache it would not: ggml's CPU flash attention splits one
query's keys among the threads from 512 keys on, so a block of one
position would change in its last bits with max_steps. Every block has
its own graph, and they share one allocator kept with the backbone.

Chunked and one-shot logits agree to float rounding, not bit for bit:
blocks use the flash attention that decode steps use. On a small
synthetic model with 64-wide heads they agree to within 5e-7 of the
largest logit on the CPU, and to within 6e-4 on Metal and CUDA, whose
flash attention rounds part of its input to half precision. On the
EN F16 checkpoint, from 256 to 8192 steps, they agree to within 1.1%
of the largest logit on CUDA, 0.15% on Metal, 0.07% on an x86 CPU and
0.6% on an M3 CPU, with the same argmax.

The tests prefill a 600-step prompt of text, audio rows across the
first block edge and a reply of tokens and frames across the second.
Cut on either side of each edge, it matches the double precision
reference and the one-shot prefill when chunked, and so do the steps
after it. A 340-step reply prefilled chunked matches feeding it step by
step. Prefills into caches of 768 and 1536 steps give the same bits,
and so does a chunked request after longer, shorter, one-shot and text
requests and in a cache another prompt filled. Both prefills reject
bad frames, NaN audio and requests past the context.
liquid-audio's ChatState keeps every turn of a conversation, and each
reply prefills all of it again. make_lfm2_chat_prompt builds that
sequence for a turn after earlier ones, as the README's multi-turn
example holds it: the system prompt once, each earlier question's audio
positions, each reply step by step with <|im_end|>\n<|im_start|>user\n
after it, then the new question. demo/chat.py, as written, adds the
system turn again on every turn, as its len(chat.text) == 1 check
always holds; this prompt keeps it once. A reply goes in as the token
ids and frames the generator yielded, the end-of-audio frame included
and a reply cut off at max_tokens kept as it is, as ChatState keeps
them, so no text is tokenized again. Without history it is the prompt
S2S already builds.

A request is to carry its earlier turns in its input artifacts: for
each turn in order, an lfm2_audio.question with the question's WAV
bytes (kind Custom) and the lfm2_audio.reply artifact that turn's
result returned. The reply artifact is kind AcousticTokens, with a
payload of little-endian int32 values: a token id per text step, and
-1 followed by one code per codebook per frame, about 470 bytes per
second of reply audio. Its meta names the format, lfm2_audio.reply/1,
and the checkpoint's language, codebook count, codebook size and
vocabulary size, which must match where it is read, and holds the step
count and whether the reply ended. It carries ids and codes only, so
it replays on any quantization of the checkpoint.

read_lfm2_conversation reads that history. It rejects artifacts it
does not know, artifacts out of turn, a question without its reply, a
WAV that does not read, and a reply whose kind, format, meta, length,
tokens or codes are not what make_lfm2_reply_artifact writes for the
checkpoint.

A conversation's prompt and max_tokens may take 8192 steps together,
and require_lfm2_conversation_room throws a CapacityError past that,
saying to leave out the oldest turns. read_lfm2_conversation counts
each reply from its meta before it reads the payload, and stops at the
first reply that leaves no room for max_tokens, so a history over the
limit is not read in full only to be turned away.
lfm2_chat_prompt_text_steps counts the prompt's steps other than the
questions' audio, so that the limit can be checked before any question
is encoded. Nothing calls these yet.

lfm2_audio_chat_test checks the prompt against ChatState's sequence for
zero, one and two earlier turns on the byte vocabulary, the text steps
of each, the artifact's bytes and its round trip, the limit, and each
rejection.
An S2S request answered one question as a new conversation. Its input
artifacts now carry the earlier turns, an lfm2_audio.question and an
lfm2_audio.reply for each, and with the new request option
return_codes the result, offline or streamed, also returns the reply
as an lfm2_audio.reply artifact to send back with the next turn. The
session keeps nothing between requests.

Each turn prefills the whole conversation again, as liquid-audio's
chat demo does: each earlier question goes through the encoder on its
own, the embeddings go into the chat prompt in order, and the earlier
replies go in as their tokens and frames. A turn with history takes
the chunked prefill. A turn without it builds the prompt it always did
and takes the one-shot prefill, so a request without artifacts or
return_codes gives the bytes it gave before.

The reply artifact holds every step the generator yielded, the
end-of-audio frame included. A reply cut off at max_tokens, and a
stream finished before its reply is over, return the steps so far with
ended=false.

A conversation's prompt and max_tokens may take 8192 steps together;
past that a turn fails with a CapacityError that says to leave out the
oldest turns, and no turn is ever dropped silently. The limit is
checked as soon as the request shows it passed: when the request is
read, on its replies and chat markup, so that a stream fails before its
audio comes, then after each earlier question is encoded, and last on
the whole prompt. A first turn is held only to the context, as before.
Earlier questions are held to lfm2_audio.max_pass_seconds like the new
one when the request is read, so a stream with a bad history fails
before its audio comes.

ASR and TTS turn away return_codes and lfm2_audio.* artifacts, and S2S
turns away any other artifact. The model spec gains return_codes.
The S2S session had no unit test that runs a reply, as the synthetic
packages had no vocoder or detokenizer. lfm2_audio_test_package.h gains
writers for both, and lfm2_audio_chat_test builds an S2S package on
them. Its backbone is random but leaves the dimensions its inputs need
to those inputs, so every reply writes "hi", speaks seven frames whose
first codes follow a fixed path, and ends on the end-of-audio frame. One
head of the last attention layer averages a mark that only the
end-of-audio frames of earlier replies carry into the logits of the
second codebook, so earlier turns change its codes.

On that package, offline: a first turn with and without return_codes,
the reply artifact's steps and meta, a second and a third turn with
the history the turns before returned, which changes the reply, each
turn again, in another order and in a fresh session, to the bit, a
reply cut at max_tokens and a turn after it, and the rejections:
artifacts of another id or out of turn, a reply of the other
checkpoint, return_codes that is not a boolean, a conversation over
8192 steps, turned away by each check from its replies to its whole
prompt, next to a turn exactly at the limit, while a first turn is not
held to it, an earlier question over lfm2_audio.max_pass_seconds, and
history or return_codes on ASR and TTS. Streamed: each turn returns
the offline turn's reply artifact, again to the bit, a stream finished
early or before its reply returns the steps so far as not ended, and a
bad history, or one whose text passes the limit, fails at
start_stream.

test_lfm2_audio_s2s gains a conversation on the published package: the
first reply's artifact holds its text and speech, a second turn
(assets/resources/a.wav) sent with the first gets another reply than
without it and the same bytes when sent again, and a third turn
replays both.
The docs said S2S answers one turn per request, as a new conversation.
A Conversations section now says how a request carries the turns
before it: the lfm2_audio.question and lfm2_audio.reply artifacts in
conversation order, the reply artifact's payload and meta, that a turn
depends on its request alone, the 8192-step limit and when a request
over it fails, how to send the history through the C API, and that the
server returns the reply artifact but does not take request artifacts
yet.

The prompt is the one ChatState holds in liquid-audio's README
multi-turn example, with the system prompt once. The section says that
demo/chat.py, as written, adds the system turn again on every turn and
that audio.cpp does not, as in ten-turn English conversations run with
liquid-audio the repeated turn left replies answering 1 of 21 questions
about earlier turns, against 7 of 21. It also gives the range checked,
10 turns with prompts of up to about 5000 steps, in which replies used
the earlier turns but not what users said about themselves and late
Japanese replies sometimes repeated an earlier one, and says to leave
out the oldest turns past it.

The request options list return_codes, the single-turn bullet leaves
the limitations, and the validation section covers the new tests.
The model spec gained the return_codes request option, so the page
is rebuilt with npm ci and npm run build in webui/native from this
branch's webui/ and model_specs/. It differs from the base's page
only in that option's entry and the bundle's hashed name. Built the
same way, the base's sources (cf124a6) give its committed page byte
for byte.
Of the errors a session throws, only CapacityError reaches a client as
a 400. Any other, a std::runtime_error or a std::invalid_argument alike,
gets 500 server_error from both server runtimes, also when the request
is at fault: an input the model does not take, or one that does not
read. The caller then looks for a server fault that is not there.
std::invalid_argument would not tell them apart, as sessions throw it
for their own faults too.

engine::runtime::InvalidRequestError, next to CapacityError, is for
those. It is a std::runtime_error, so the CLI and the C API report it
as before, and both server runtimes answer it with
400 invalid_request_error. CapacityError stays as it is: there the
request is fine, only too big for the device. Nothing throws the new
error yet.

The parallel lifecycle test gets a session that turns its request away
and checks the 400 on /v1/tasks/run, offline and streaming, and on
/v1/tasks/stream, with the lease released each time, and that another
failure still leaves the handler for the transport's 500.
Through the server, history that S2S turns away (out of turn, another
format, checkpoint or kind, bad meta or payload, a question that does
not read as WAV), a return_codes that is neither true nor false, and
history or return_codes sent to ASR or TTS came back as
500 server_error. They are the caller's to fix, so they now throw
InvalidRequestError, which both server runtimes answer with 400.

So does reject_options, which turns return_codes away on ASR and TTS,
and with it the other options a task does not take, a 500 on main too.
A reply step about to be written that does not fit the checkpoint stays
a std::runtime_error, a fault of the session, so check_step takes the
error type. The conversation limit stays a CapacityError. lfm2_audio's
other checks, and the option parsers every family shares, are
unchanged.

lfm2_audio_chat_test now requires InvalidRequestError for each request
it turns away, and an error of another type where a prompt or a reply
artifact is built from bad steps. The docs say which errors are a 400.
@0xShug0

0xShug0 commented Oct 8, 2026

Copy link
Copy Markdown
Owner

OOTO now will check later today...

@0xShug0
0xShug0 merged commit c0d84e1 into 0xShug0:main Oct 8, 2026
11 checks passed
@0xShug0

0xShug0 commented Oct 8, 2026

Copy link
Copy Markdown
Owner

@ykhrustalev PR merged! One thing can be addressed in a follow-up PR: use #834’s packed-buffer-aware validator in the new BlockGraph, consistent with the existing prefill and decode graphs. But no actual failure from this omission was found.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants