server: support inp embd to generate next token - #26313
Conversation
|
/bot review |
Automated code reviewStatic review of PR #26313 (server: support inp embd to generate next token). Blocking(point 1) This is reachable on the normal path: ~server_batch() {
if (batch.pos != nullptr) {
if (batch.token == nullptr) { // we are in embd-rendered state
batch.token = tokens_ptr;
batch.embd = nullptr;
}
llama_batch_free(batch);
}
}(point 2) Will slow the review(point 3) (point 4) When a generating slot runs in embd mode, the prompt-processing loop silently skips prompt-token slots via (point 5) The new Nits(point 6) The member (point 7) Overall: the implementation approach (reuse one batch, swap This review was generated automatically by pi coding agent using |
hmm yeah that need to be fixed |
|
@ggerganov could you take a quick look? thanks! |
* server: support embd for sampled token * fix ~server_batch()
llama_model_n_embd_inp() takes const llama_model *, but upstream ggml-org#26313 passed the llama_model_ptr directly; the rebased tree does not compile without .get(). Assisted-by: pi
* server: support embd for sampled token * fix ~server_batch()
* server: support embd for sampled token * fix ~server_batch()
* server: support embd for sampled token * fix ~server_batch()
* server: support embd for sampled token * fix ~server_batch()
* server: support embd for sampled token * fix ~server_batch()
Overview
Part of #26254
When generating the next token, TTS backbone model outputs both the sampled token (for semantic code) and hidden state - this is similar to MTP where the MTP head takes both the sampled and last embd
The only difference is that MTP gives back draft tokens to tgt model, while TTS model gives back one embd row (= sum of generated codes), so this PR adds support for that case
Note that we cannot reuse the same code path from mtmd-helper (for decode embd), because that doesn't contain the sampling step
Requirements