Repository navigation
Conversation
Convert Responses API input_image content in function_call_output to Chat Completions image_url content. This allows multimodal tool results, such as Codex view_image output, to pass through the existing multimodal pipeline. Add regression tests for text/image tool outputs and invalid image content.
|
Hi @kossum, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
/bot review |
Automated code reviewReview of PR #28847 - server: support input_image in Responses tool outputsI checked the conversion logic end-to-end: the produced Chat Completions Minor points: (point 1) Missing type check on (point 2) The (point 3) Nit: the two branches are asymmetric - Style and conventions look clean: ASCII only, concise comments, error message wording matches the surrounding code, and the tests follow the existing "Expected exception" pattern already used in this file (e.g. line 1658). This review was generated automatically by pi coding agent using |
a7f76f8 to
e3d2144
Compare
The function_call_output branches now check field types with exists_and_is_string instead of a bare contains(): - input_image: image_url must be a string, so a non-string value (e.g. a Chat-Completions-style object) fails fast with a clear message here instead of surfacing later in the media pipeline - input_text: text must be present and a string, consistent with the user-message path
eda8c92 to
bd1d410
Compare
|
I've addressed the automated review feedback in this commit.
The new test code was generated by Qwen3.8-27B; I reviewed the logic, adjusted it, and verified locally that all tests pass (including the new negative cases). |
|
duplicated of #22575 |
Overview
Support
input_imagecontent in Responses APIfunction_call_output.Codex's
view_imagetool returns image content as aninput_imageblock inside
function_call_output.output. llama-server currentlyonly accepts
input_textblocks there and rejects image tool results.This converts
input_imagetool output to the existing ChatCompletions
image_urlrepresentation so it can pass through theexisting multimodal pipeline.
Additional information
Added unit tests covering:
input_textarray outputinput_textandinput_imagetool outputimage_urlRelated to #23890
Requirements