Skip to content

server : fix tool calls being dropped from trailing assistant message - #27626

Merged
aldehir merged 2 commits into
ggml-org:masterfrom
kyo-zzz:fix/continuation-tool-calls
Aug 25, 2026
Merged

aldehir merged 2 commits into
ggml-org:masterfrom
kyo-zzz:fix/continuation-tool-calls

Conversation

@kyo-zzz

@kyo-zzz kyo-zzz commented Aug 23, 2026 •

Copy link
Copy Markdown

Fixes #27588

Overview

The bug

--prefill-assistant + trailing assistant(tool_calls) = silent context corruption. Request switches to continuation mode, add_generation_prompt forced off, prompt tail rebuilt without the tool calls. Model continues from an assistant turn where those calls don't exist. Agent clients replaying such histories get hosed.

Why

Three spots:

  1. tools/server/server-common.cpp — heuristic triggers on any trailing assistant msg
  2. common/chat.cpp — final msg popped from render list
  3. common/chat-auto-parser-generator.cpp — continuation suffix skips tool_calls entirely

Fix

  • Auto-continuation now skips trailing assistant msgs that have tool calls → they render normally, add_generation_prompt is honored
  • Explicit continue_final_message throws instead of silently corrupting (consistent with the existing conflict error)

Tested

Regression tests in tests/test-chat.cpp (Qwen3): normal render keeps get_weather, explicit continuation is rejected, server parse path keeps both the tool call and gen prompt.

Additional information

Thanks to @PieBru for providing a GPU-free reproduction of the /apply-template issue in #27588. This PR follows the "reject loudly" option discussed there.

Built and tested locally on Windows 11 with MSVC 2022 (static build).

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES — I used AI to assist with root-cause analysis, code changes, and test development. I reviewed every line, built and ran the tests locally, and take full responsibility for the change.

@kyo-zzz
kyo-zzz requested review from a team and pwilkin as code owners August 23, 2026 20:25
@github-actions github-actions Bot added testing Everything test related server labels Aug 23, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Aug 23, 2026

Copy link
Copy Markdown

Hi @kyo-zzz, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • AI-generated content: While code is allowed to be generated by AI, please write the PR description and commit messages on your own without the help of AI.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

…tant

Last assistant carries tool_calls + --prefill-assistant is on → request
flips into continuation mode, add_generation_prompt forced off, tail
rebuilt from reasoning_content + content only. Tool calls just vanish.

- Auto-continuation now skips trailing assistant msgs that have tool calls
- continue_final_message on those throws a clear error instead of
  silently corrupting the prompt
- Regression tests included, red before / green after

Fixes ggml-org#27588

Developed with AI assistance, disclosed per the contribution policy.
@kyo-zzz
kyo-zzz force-pushed the fix/continuation-tool-calls branch from 6954c2a to 9d87e6d Compare August 23, 2026 21:08
Comment thread tools/server/server-common.cpp
Comment thread common/chat.cpp Outdated
Comment thread tools/server/server-common.cpp Outdated
Move validation into oaicompat_chat_params_parse (next to the existing
two-or-more-assistant check) and remove it from common_chat_templates_apply,
which has no precedent for validation. Drop the regression tests.

Per review: --prefill-assistant with a trailing assistant message
containing tool calls is not supported and should fail loudly.
@kyo-zzz
kyo-zzz force-pushed the fix/continuation-tool-calls branch from cd2e017 to aeba8f1 Compare August 24, 2026 06:35
@zhong285

zhong285 commented Aug 25, 2026 •

Copy link
Copy Markdown

I independently reproduced this issue on an older affected llama.cpp build, but had not posted the results here previously. I have now also verified the current PR head (aeba8f159d185afd1ce5384c4c161842ccb3eb89) against the same agent-style history boundary.

Environment

Affected build:

  • b1-a94d563ed

PR-head build:

  • b339-aeba8f1
  • source: aeba8f159d185afd1ce5384c4c161842ccb3eb89

Common runtime profile:

  • Qwen3.6-35B-A3B Q4_K_M
  • Vulkan
  • n-gpu-layers=999
  • threads=4
  • default assistant prefill enabled
  • reasoning disabled
  • GGML_NATIVE=OFF

The replay fixture used a real assistant(tool_calls) object originally produced by the model.

Results

1. Trailing assistant(tool_calls)

Affected build:

  • /apply-template returned HTTP 200
  • the trailing tool call disappeared from the rendered prompt
  • no error was reported

PR head:

  • the same history returned HTTP 400
  • structured error:

Cannot continue an assistant message that contains tool calls.

Result: the previously observed silent context corruption is replaced by an explicit fail-closed rejection.

2. Completed synchronous tool round

For:

assistant(tool_calls) -> tool(result)

the PR head returned HTTP 200 and preserved the function call, argument, tool result, ordering, and next assistant generation prompt.

Result: no regression observed for the completed synchronous tool round.

3. Explicit continuation

Using the same trailing assistant(tool_calls) history with:

  • continue_final_message=true
  • add_generation_prompt=false

the PR head returned HTTP 400 with the same structured rejection.

Result: explicit continuation also fails closed.

Before / after

Case Affected build PR head
trailing assistant(tool_calls) HTTP 200 + silent loss HTTP 400 fail-closed
completed assistant(tool_calls) -> tool(result) preserved preserved
explicit continuation not separately qualified HTTP 400 fail-closed

For this tested Qwen3.6/Vulkan environment, the current PR head removes the silent tool-call history corruption at the affected boundary while preserving the completed synchronous tool round.

This is a focused independent verification of this specific runtime/history boundary, not a claim of universal coverage across models or backends.

@aldehir
aldehir merged commit 1729ed5 into ggml-org:master Aug 25, 2026
23 of 28 checks passed
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
…g#27626)

* server: fix tool calls getting silently stripped with --prefill-assistant

Last assistant carries tool_calls + --prefill-assistant is on → request
flips into continuation mode, add_generation_prompt forced off, tail
rebuilt from reasoning_content + content only. Tool calls just vanish.

- Auto-continuation now skips trailing assistant msgs that have tool calls
- continue_final_message on those throws a clear error instead of
  silently corrupting the prompt
- Regression tests included, red before / green after

Fixes ggml-org#27588

Developed with AI assistance, disclosed per the contribution policy.

* server : address review: fail on prefill-assistant + trailing tool_calls

Move validation into oaicompat_chat_params_parse (next to the existing
two-or-more-assistant check) and remove it from common_chat_templates_apply,
which has no precedent for validation. Drop the regression tests.

Per review: --prefill-assistant with a trailing assistant message
containing tool calls is not supported and should fail loudly.
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…g#27626)

* server: fix tool calls getting silently stripped with --prefill-assistant

Last assistant carries tool_calls + --prefill-assistant is on → request
flips into continuation mode, add_generation_prompt forced off, tail
rebuilt from reasoning_content + content only. Tool calls just vanish.

- Auto-continuation now skips trailing assistant msgs that have tool calls
- continue_final_message on those throws a clear error instead of
  silently corrupting the prompt
- Regression tests included, red before / green after

Fixes ggml-org#27588

Developed with AI assistance, disclosed per the contribution policy.

* server : address review: fail on prefill-assistant + trailing tool_calls

Move validation into oaicompat_chat_params_parse (next to the existing
two-or-more-assistant check) and remove it from common_chat_templates_apply,
which has no precedent for validation. Drop the regression tests.

Per review: --prefill-assistant with a trailing assistant message
containing tool calls is not supported and should fail loudly.
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…g#27626)

* server: fix tool calls getting silently stripped with --prefill-assistant

Last assistant carries tool_calls + --prefill-assistant is on → request
flips into continuation mode, add_generation_prompt forced off, tail
rebuilt from reasoning_content + content only. Tool calls just vanish.

- Auto-continuation now skips trailing assistant msgs that have tool calls
- continue_final_message on those throws a clear error instead of
  silently corrupting the prompt
- Regression tests included, red before / green after

Fixes ggml-org#27588

Developed with AI assistance, disclosed per the contribution policy.

* server : address review: fail on prefill-assistant + trailing tool_calls

Move validation into oaicompat_chat_params_parse (next to the existing
two-or-more-assistant check) and remove it from common_chat_templates_apply,
which has no precedent for validation. Drop the regression tests.

Per review: --prefill-assistant with a trailing assistant message
containing tool calls is not supported and should fail loudly.
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
…g#27626)

* server: fix tool calls getting silently stripped with --prefill-assistant

Last assistant carries tool_calls + --prefill-assistant is on → request
flips into continuation mode, add_generation_prompt forced off, tail
rebuilt from reasoning_content + content only. Tool calls just vanish.

- Auto-continuation now skips trailing assistant msgs that have tool calls
- continue_final_message on those throws a clear error instead of
  silently corrupting the prompt
- Regression tests included, red before / green after

Fixes ggml-org#27588

Developed with AI assistance, disclosed per the contribution policy.

* server : address review: fail on prefill-assistant + trailing tool_calls

Move validation into oaicompat_chat_params_parse (next to the existing
two-or-more-assistant check) and remove it from common_chat_templates_apply,
which has no precedent for validation. Drop the regression tests.

Per review: --prefill-assistant with a trailing assistant message
containing tool calls is not supported and should fail loudly.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

server testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Trailing assistant message with tool_calls: calls silently dropped from rendered prompt (auto-prefill/continuation path)

4 participants