Skip to content

server: support input_image in function_call_output (#20663) - #22575

Merged
ngxson merged 3 commits into
ggml-org:masterfrom
Empressia:support-input_image-in-function_call_output
Sep 22, 2026
Merged

ngxson merged 3 commits into
ggml-org:masterfrom
Empressia:support-input_image-in-function_call_output

Conversation

@Empressia

Copy link
Copy Markdown
Contributor

Overview

enables multimodal input from tool outputs by supporting the processing of "input_image" in "function_call_output".

Additional information

related to #20663

tested using the Gemma 4 model.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES (used for translating documentation and seeking advice on development tooling. implementation and testing were performed manually.)

@Empressia
Empressia requested a review from a team as a code owner May 1, 2026 06:57
@cereblab

cereblab commented Jun 9, 2026

Copy link
Copy Markdown

Tested — works. 👍

Use case: the OpenAI Codex CLI driving a local llama-server over the Responses API (wire_api=responses). Codex's built-in view_image tool returns the image as an input_image inside a function_call_output, which previously hit 400 "Output of tool call should be 'Input text'" (#20663) and aborted the turn (poisoning the conversation history, since the rejected item gets re-sent on every replay).

Built llama-server with this patch and sent a /v1/responses request containing an image in a function_call_output (gemma-4 + mmproj):

  • Before: HTTP 400 — Output of tool call should be 'Input text'
  • After: HTTP 200, and the model correctly described the image.

This confirms the image flows through the existing media pipeline — it lands as a tool message whose image_url content is handled by the standard message loop in server-common.cpp (no role-gating), exactly as this PR intends.

This is currently the only blocker for vision-capable agentic tools (Codex, etc.) running against local models on the Responses API. Would be great to see it merged. 🙏

@Nico3012

Copy link
Copy Markdown

Any updates on merging this?
@ngxson @ggerganov would appreciate any feedback when you have a moment.

@RufiS

RufiS commented Jun 26, 2026 •

Copy link
Copy Markdown

I hit this same issue with Codex image viewing.

My setup:

  • llama.cpp
  • Qwen3.6-27B dense / qwen35
  • mmproj enabled
  • Codex using the Responses API image-viewing path
  • server running with --jinja, --reasoning on, --spec-type draft-mtp, --tools all, and a local mmproj

The failure was:

400 "Output of tool call should be 'Input text'"

I have a private-fork workaround that has been stable for my local setup, but that workaround was AI-assisted and I cannot responsibly submit or maintain it upstream because I cannot explain every line to the standard expected by this project.

I am adding this only as reproduction / validation data: this PR appears to target a real blocker for Codex image workflows against local vision-capable llama.cpp models.

@mebassett

Copy link
Copy Markdown

Just chiming in that I've hit the same bug and fixed it by using a build based on merging this PR into my local master

@rpc180

rpc180 commented Aug 21, 2026 •

Copy link
Copy Markdown

Hitting the same 400 and confirm its still present in a llama.cpp build from last week. Merging this PR with my local master also resolved the problem.

• llama.cpp: ghcr.io/ggml-org/llama.cpp:server-cuda, build fingerprint b10450 → commit
ece963f (2026-08-15)
• Model: Qwen 3.8 27B (UD Q6_K_XL GGUF) + mmproj-F16.gguf, --jinja, --ctx-size 131072,
--n-gpu-layers all, flash-attn on
OpenAI Codex (v0.149.0)
• Client: Codex CLI against POST /v1/responses
• Trigger: model invokes the view_image tool → the function_call_output item carries output:
[{type: "input_image", image_url: "data:image/png;base64,..."}] → server throws the 400, Codex
aborts the session. Every subsequent turn replays the poisoned history, so the session is
unusable until reset.
• Server log: W srv operator(): got exception: {"error":{"code":400,"message":"Output of tool
call should be 'Input text'","type":"invalid_request_error"}}

Key diagnostic that narrows it to the converter, not the multimodal path: a direct curl to
/v1/chat/completions with the same image as an image_url content part succeeds — the
model correctly describes the image (2736 prompt tokens, finish_reason: stop). So image
encoding/mmproj is healthy; the failure is isolated to the Responses→ChatCompletions
function_call_output conversion in server-chat.cpp, which is exactly what this PR changes.

Also confirms #20663 is not model-version-specific (that thread reported Qwen 3.5/3.6; this is    
3.8 27B).                                                                                  

@lee-b

lee-b commented Sep 4, 2026

Copy link
Copy Markdown

@ggerganov approved above and clean. Could you please merge? Neither codex (from above comments) nor pi (from my own experience with this issue) can view images right now.

@ServeurpersoCom ServeurpersoCom left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, the conversion is correct and lands the image in the standard media path: the url shape is what handle_media expects, and the allow_image gate still applies since the converted part goes through the regular content loop. Merges cleanly on current master, and nothing here widens the remote URL surface that user messages already reach.

Two nits, neither blocking:

The two if statements added in the dispatch are missing a space before the parenthesis.

The type is read back four times through at("type"). A local for it plus a final else gives one lookup and, more usefully, a single place listing the accepted types instead of duplicating them between the guard and the dispatch.

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

You need another approval, the one above is from a contributor without write access. cc @ngxson

@Empressia

Copy link
Copy Markdown
Contributor Author

I adjusted the if statement spacing and stored type in a local variable to avoid repeated lookups.

@lee-b

lee-b commented Sep 9, 2026

Copy link
Copy Markdown

You need another approval, the one above is from a contributor without write access. cc @ngxson

If this is referring to me, I was not claiming to have approved it, I was saying that you had approved it, and asking for another approval + merge :D

@gpez-git

Copy link
Copy Markdown

Would be great to get another approver - this seems to be a substantial issue for those using vision capable models. I can confirm OPM is affected as well.

@ngxson
ngxson merged commit 4098fdc into ggml-org:master Sep 22, 2026
1 check passed
LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
…gml-org#22575)

* server: support input_image in function_call_output (ggml-org#20663)

* server: fix if statement spacing

* server: avoid repeated type lookup
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…gml-org#22575)

* server: support input_image in function_call_output (ggml-org#20663)

* server: fix if statement spacing

* server: avoid repeated type lookup
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…gml-org#22575)

* server: support input_image in function_call_output (ggml-org#20663)

* server: fix if statement spacing

* server: avoid repeated type lookup

(cherry picked from commit 4098fdc)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants