server : fix tool calls being dropped from trailing assistant message - #27626
Conversation
|
Hi @kyo-zzz, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
…tant Last assistant carries tool_calls + --prefill-assistant is on → request flips into continuation mode, add_generation_prompt forced off, tail rebuilt from reasoning_content + content only. Tool calls just vanish. - Auto-continuation now skips trailing assistant msgs that have tool calls - continue_final_message on those throws a clear error instead of silently corrupting the prompt - Regression tests included, red before / green after Fixes ggml-org#27588 Developed with AI assistance, disclosed per the contribution policy.
6954c2a to
9d87e6d
Compare
Move validation into oaicompat_chat_params_parse (next to the existing two-or-more-assistant check) and remove it from common_chat_templates_apply, which has no precedent for validation. Drop the regression tests. Per review: --prefill-assistant with a trailing assistant message containing tool calls is not supported and should fail loudly.
cd2e017 to
aeba8f1
Compare
|
I independently reproduced this issue on an older affected llama.cpp build, but had not posted the results here previously. I have now also verified the current PR head ( EnvironmentAffected build:
PR-head build:
Common runtime profile:
The replay fixture used a real Results1. Trailing Affected build:
PR head:
Result: the previously observed silent context corruption is replaced by an explicit fail-closed rejection. 2. Completed synchronous tool round For:
the PR head returned HTTP 200 and preserved the function call, argument, tool result, ordering, and next assistant generation prompt. Result: no regression observed for the completed synchronous tool round. 3. Explicit continuation Using the same trailing
the PR head returned HTTP 400 with the same structured rejection. Result: explicit continuation also fails closed. Before / after
For this tested Qwen3.6/Vulkan environment, the current PR head removes the silent tool-call history corruption at the affected boundary while preserving the completed synchronous tool round. This is a focused independent verification of this specific runtime/history boundary, not a claim of universal coverage across models or backends. |
…g#27626) * server: fix tool calls getting silently stripped with --prefill-assistant Last assistant carries tool_calls + --prefill-assistant is on → request flips into continuation mode, add_generation_prompt forced off, tail rebuilt from reasoning_content + content only. Tool calls just vanish. - Auto-continuation now skips trailing assistant msgs that have tool calls - continue_final_message on those throws a clear error instead of silently corrupting the prompt - Regression tests included, red before / green after Fixes ggml-org#27588 Developed with AI assistance, disclosed per the contribution policy. * server : address review: fail on prefill-assistant + trailing tool_calls Move validation into oaicompat_chat_params_parse (next to the existing two-or-more-assistant check) and remove it from common_chat_templates_apply, which has no precedent for validation. Drop the regression tests. Per review: --prefill-assistant with a trailing assistant message containing tool calls is not supported and should fail loudly.
…g#27626) * server: fix tool calls getting silently stripped with --prefill-assistant Last assistant carries tool_calls + --prefill-assistant is on → request flips into continuation mode, add_generation_prompt forced off, tail rebuilt from reasoning_content + content only. Tool calls just vanish. - Auto-continuation now skips trailing assistant msgs that have tool calls - continue_final_message on those throws a clear error instead of silently corrupting the prompt - Regression tests included, red before / green after Fixes ggml-org#27588 Developed with AI assistance, disclosed per the contribution policy. * server : address review: fail on prefill-assistant + trailing tool_calls Move validation into oaicompat_chat_params_parse (next to the existing two-or-more-assistant check) and remove it from common_chat_templates_apply, which has no precedent for validation. Drop the regression tests. Per review: --prefill-assistant with a trailing assistant message containing tool calls is not supported and should fail loudly.
…g#27626) * server: fix tool calls getting silently stripped with --prefill-assistant Last assistant carries tool_calls + --prefill-assistant is on → request flips into continuation mode, add_generation_prompt forced off, tail rebuilt from reasoning_content + content only. Tool calls just vanish. - Auto-continuation now skips trailing assistant msgs that have tool calls - continue_final_message on those throws a clear error instead of silently corrupting the prompt - Regression tests included, red before / green after Fixes ggml-org#27588 Developed with AI assistance, disclosed per the contribution policy. * server : address review: fail on prefill-assistant + trailing tool_calls Move validation into oaicompat_chat_params_parse (next to the existing two-or-more-assistant check) and remove it from common_chat_templates_apply, which has no precedent for validation. Drop the regression tests. Per review: --prefill-assistant with a trailing assistant message containing tool calls is not supported and should fail loudly.
…g#27626) * server: fix tool calls getting silently stripped with --prefill-assistant Last assistant carries tool_calls + --prefill-assistant is on → request flips into continuation mode, add_generation_prompt forced off, tail rebuilt from reasoning_content + content only. Tool calls just vanish. - Auto-continuation now skips trailing assistant msgs that have tool calls - continue_final_message on those throws a clear error instead of silently corrupting the prompt - Regression tests included, red before / green after Fixes ggml-org#27588 Developed with AI assistance, disclosed per the contribution policy. * server : address review: fail on prefill-assistant + trailing tool_calls Move validation into oaicompat_chat_params_parse (next to the existing two-or-more-assistant check) and remove it from common_chat_templates_apply, which has no precedent for validation. Drop the regression tests. Per review: --prefill-assistant with a trailing assistant message containing tool calls is not supported and should fail loudly.
Fixes #27588
Overview
The bug
--prefill-assistant+ trailingassistant(tool_calls)= silent context corruption. Request switches to continuation mode,add_generation_promptforced off, prompt tail rebuilt without the tool calls. Model continues from an assistant turn where those calls don't exist. Agent clients replaying such histories get hosed.Why
Three spots:
tools/server/server-common.cpp— heuristic triggers on any trailing assistant msgcommon/chat.cpp— final msg popped from render listcommon/chat-auto-parser-generator.cpp— continuation suffix skipstool_callsentirelyFix
add_generation_promptis honoredcontinue_final_messagethrows instead of silently corrupting (consistent with the existing conflict error)Tested
Regression tests in
tests/test-chat.cpp(Qwen3): normal render keepsget_weather, explicit continuation is rejected, server parse path keeps both the tool call and gen prompt.Additional information
Thanks to @PieBru for providing a GPU-free reproduction of the
/apply-templateissue in #27588. This PR follows the "reject loudly" option discussed there.Built and tested locally on Windows 11 with MSVC 2022 (static build).
Requirements