Skip to content

server: support input_image in Responses tool outputs - #28847

Closed
kossum wants to merge 2 commits into
ggml-org:masterfrom
kossum:fix/responses-tool-output-image
Closed

kossum wants to merge 2 commits into
ggml-org:masterfrom
kossum:fix/responses-tool-output-image

Conversation

@kossum

@kossum kossum commented Sep 13, 2026 •

Copy link
Copy Markdown

Overview

Support input_image content in Responses API function_call_output.

Codex's view_image tool returns image content as an input_image
block inside function_call_output.output. llama-server currently
only accepts input_text blocks there and rejects image tool results.

This converts input_image tool output to the existing Chat
Completions image_url representation so it can pass through the
existing multimodal pipeline.

Additional information

Added unit tests covering:

  • existing input_text array output
  • mixed input_text and input_image tool output
  • missing image_url
  • unsupported output content types

Related to #23890

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES — AI was used to assist with the investigation and test implementation. Qwen3.8-27B generated the initial regression test code, which I reviewed, modified, and tested manually. ChatGPT was used to discuss the issue and implementation approach.

Convert Responses API input_image content in function_call_output to Chat Completions image_url content.

This allows multimodal tool results, such as Codex view_image output, to pass through the existing multimodal pipeline.

Add regression tests for text/image tool outputs and invalid image content.
@kossum
kossum requested review from a team and pwilkin as code owners September 13, 2026 13:46
@github-actions github-actions Bot added testing Everything test related server labels Sep 13, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 13, 2026

Copy link
Copy Markdown

Hi @kossum, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Sep 13, 2026
@github-actions
github-actions Bot marked this pull request as draft September 13, 2026 13:51
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Sep 13, 2026
@kossum
kossum marked this pull request as ready for review September 13, 2026 14:12
@pwilkin

pwilkin commented Sep 13, 2026

Copy link
Copy Markdown
Member

/bot review

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Automated code review

Review of PR #28847 - server: support input_image in Responses tool outputs

I checked the conversion logic end-to-end: the produced Chat Completions image_url part flows correctly through oaicompat_chat_params_parse (tools/server/server-common.cpp:1238, which accepts array content on any role) and common_chat_msgs_parse_oaicompat, so the change is functionally sound. The scope is good too: single-purpose, reuses the existing user-message input_image handling pattern, tests added to the existing test_convert_responses_to_chatcmpl case rather than a new file, and AI usage is disclosed. No blocking issues found.

Minor points:

(point 1) Missing type check on image_url value - tools/server/server-chat.cpp, server_chat_convert_responses_to_chatcmpl, the input_image branch of the tool-output loop. The code checks contains("image_url") but not that it is a string. If a client sends a Chat-Completions-style object ("image_url": {"url": ...}) or another type, it gets wrapped as {"url": <object>} and the failure surfaces much later in handle_media as a confusing "Invalid base64 value". An exists_and_is_string-style check (the helper is already defined in this function) with a clear error would match how the sibling branches validate text. Note the user-message input_image path (line ~82) has the same gap, so fixing both in one go would be natural.

(point 2) The input_text branch of the same loop never verifies that text is present/a string, unlike the user-message path which throws "'Input text' requires 'text'". A malformed {"type": "input_text"} part currently fails much later with "key 'text' not found" from common_chat_msgs_parse_oaicompat. This is pre-existing behavior, but since the loop is being reworked here, adding the check would make the errors consistent across the converter.

(point 3) Nit: the two branches are asymmetric - input_text is converted by mutating the original part (extra fields like annotations pass through), while input_image rebuilds the part from scratch (dropping extra fields such as detail). Behaviorally fine since Chat Completions consumers ignore the extras, just something to be aware of.

Style and conventions look clean: ASCII only, concise comments, error message wording matches the surrounding code, and the tests follow the existing "Expected exception" pattern already used in this file (e.g. line 1658).

This review was generated automatically by pi coding agent using zai-org/GLM-5.3. It may contain mistakes. Maintainers make the final call.

@kossum
kossum force-pushed the fix/responses-tool-output-image branch from a7f76f8 to e3d2144 Compare September 14, 2026 00:56
The function_call_output branches now check field types with
exists_and_is_string instead of a bare contains():

- input_image: image_url must be a string, so a non-string value
  (e.g. a Chat-Completions-style object) fails fast with a clear
  message here instead of surfacing later in the media pipeline
- input_text: text must be present and a string, consistent with the
  user-message path
@kossum
kossum force-pushed the fix/responses-tool-output-image branch from eda8c92 to bd1d410 Compare September 14, 2026 01:26
@kossum

kossum commented Sep 14, 2026

Copy link
Copy Markdown
Author

I've addressed the automated review feedback in this commit.

  • image_url in input_image tool output is now validated as a string (via exists_and_is_string), so a non-string value now fails fast with a clear message here instead of surfacing later in the media pipeline.
  • input_text tool output now requires text to be a string, consistent with the user-message path.

The new test code was generated by Qwen3.8-27B; I reviewed the logic, adjusted it, and verified locally that all tests pass (including the new negative cases).

@ngxson

ngxson commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

duplicated of #22575

@ngxson ngxson closed this Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

server testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants