Conversation
mapReasoningEffortToThinkingLevel rejected reasoning_effort "high" for any Gemini 3 model whose name does not contain "flash", so a Chat Completions request against gemini-3-pro with reasoning_effort: high failed with an invalid request body error before reaching Vertex. The guard was carried over from the "none" case, where it is right because only Flash supports the minimal thinking level. Google's thinking docs list high as a supported level for every Gemini 3 model and as the default for the Pro models, and the function already maps "medium" on Pro to ThinkingLevelHigh. Map "high" to ThinkingLevelHigh for all models and cover the Pro case in both the unit table and the generation config table. AI assistance: Claude was used to help find the mismatch and draft the change; it was reviewed and tested locally. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: AlSh007 <alok44170@gmail.com>
✅ Deploy Preview for theagentrouter canceled.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
mapReasoningEffortToThinkingLevelrejectsreasoning_effort: "high"for any Gemini 3 model whose name does not containflash. A Chat Completions request againstgemini-3-prowithreasoning_effort: "high"therefore fails withinvalid reasoning effort: ... reasoning effort 'high' is only supported for Gemini Flash modelsbefore it reaches Vertex, while the same request withmediumsucceeds and is sent asthinking_level: high.The guard looks carried over from the
nonecase, where it is correct because only Flash supports theminimalthinking level. Forhighit is not: Google's thinking docs [1] listhighas a supported level for every Gemini 3 model and as the default level forgemini-3-pro-previewandgemini-3.1-pro-preview, the Vertex OpenAI compatibility guide [2] does not restrict it, and the function's own doc comment already says"high" → ThinkingLevelHighwithout a model restriction.This change maps
hightoThinkingLevelHighfor all models and adds the Pro case toTestMapReasoningEffortToThinkingLeveland to theopenAIReqToGeminiGenerationConfigtable. The new rows fail on main with the error above and pass with the fix.go test ./internal/translator/passes.AI usage: I used Claude Code [3] to help find the mismatch and draft the change; I reviewed the code and ran the tests locally.
Related Issues/PRs (if applicable)
The guard was introduced in #1844, whose
hightest only coversgemini-3-flash.1: https://ai.google.dev/gemini-api/docs/thinking
2: https://docs.cloud.google.com/vertex-ai/generative-ai/docs/start/get-started-with-gemini-3#openai-example
3: https://claude.com/claude-code