Repository navigation
Conversation
|
Tested — works. 👍 Use case: the OpenAI Codex CLI driving a local Built
This confirms the image flows through the existing media pipeline — it lands as a This is currently the only blocker for vision-capable agentic tools (Codex, etc.) running against local models on the Responses API. Would be great to see it merged. 🙏 |
|
Any updates on merging this? |
|
I hit this same issue with Codex image viewing. My setup:
The failure was: I have a private-fork workaround that has been stable for my local setup, but that workaround was AI-assisted and I cannot responsibly submit or maintain it upstream because I cannot explain every line to the standard expected by this project. I am adding this only as reproduction / validation data: this PR appears to target a real blocker for Codex image workflows against local vision-capable |
|
Just chiming in that I've hit the same bug and fixed it by using a build based on merging this PR into my local master |
|
Hitting the same 400 and confirm its still present in a llama.cpp build from last week. Merging this PR with my local master also resolved the problem. • llama.cpp: ghcr.io/ggml-org/llama.cpp:server-cuda, build fingerprint b10450 → commit Key diagnostic that narrows it to the converter, not the multimodal path: a direct curl to |
|
@ggerganov approved above and clean. Could you please merge? Neither codex (from above comments) nor pi (from my own experience with this issue) can view images right now. |
ServeurpersoCom
left a comment
There was a problem hiding this comment.
LGTM, the conversion is correct and lands the image in the standard media path: the url shape is what handle_media expects, and the allow_image gate still applies since the converted part goes through the regular content loop. Merges cleanly on current master, and nothing here widens the remote URL surface that user messages already reach.
Two nits, neither blocking:
The two if statements added in the dispatch are missing a space before the parenthesis.
The type is read back four times through at("type"). A local for it plus a final else gives one lookup and, more usefully, a single place listing the accepted types instead of duplicating them between the guard and the dispatch.
|
You need another approval, the one above is from a contributor without write access. cc @ngxson |
|
I adjusted the if statement spacing and stored type in a local variable to avoid repeated lookups. |
If this is referring to me, I was not claiming to have approved it, I was saying that you had approved it, and asking for another approval + merge :D |
|
Would be great to get another approver - this seems to be a substantial issue for those using vision capable models. I can confirm OPM is affected as well. |
…gml-org#22575) * server: support input_image in function_call_output (ggml-org#20663) * server: fix if statement spacing * server: avoid repeated type lookup
…gml-org#22575) * server: support input_image in function_call_output (ggml-org#20663) * server: fix if statement spacing * server: avoid repeated type lookup
…gml-org#22575) * server: support input_image in function_call_output (ggml-org#20663) * server: fix if statement spacing * server: avoid repeated type lookup (cherry picked from commit 4098fdc)
Overview
enables multimodal input from tool outputs by supporting the processing of "input_image" in "function_call_output".
Additional information
related to #20663
tested using the Gemma 4 model.
Requirements