fix(groq): replace retired models with openai/gpt-oss-120b - #161
Conversation
Groq retired qwen/qwen3-32b (validation, review; 2026-07-17) and llama-3.3-70b-versatile (generation, autofix; 2026-08-16), so every Groq call failed with 404 model_not_found. - All stages default to openai/gpt-oss-120b (GROQ_MODEL still overrides). - gpt-oss is a reasoning model: new optional <stage>_reasoning_effort key (low|medium|high), validated by loadLLMConfig(), forwarded by all four entrypoints (incl. loadConfigFromEnv) and sent by groq_client only when set, since non-reasoning models reject it. Set to low. - autofix_max_input_tokens 7400 -> 3400 to fit the 8K free-tier TPM (system + input + 4096 output ~= 7956); a test asserts the sum. - auto_fix_pr knows the 131,072-token context window. - ADR-0024; README, AGENTS, code-generation and runbook updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GzuUtVET9ZuK7cbx2LRUZM
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Code review —
|
| Area | Status | Notes |
|---|---|---|
| Test coverage | ✅ | validate_issue/generate_issue_change forwarding, the global fallback key, Anthropic ignoring the param.reasoning_effort_wiring.test.mjs (4 E2E tests at the HTTP boundary: model, max_tokens, reasoning_effort, off, override). +7 tests in config.test.mjs, +1 in anthropic_client.test.mjs. 723 → 735 → 804 after merging main. |
| Documentation coverage | ✅ | AGENTS.md:25, runbook row for 400, <stage>_reasoning_effort key in the config table.code-generation.md (<stage>, _temperature, _max_tokens, _reasoning_effort). GROQ_REASONING_EFFORT added to the env matrix. |
| ADR created | ✅ | |
| Architecture diagram | ✅ N/A | The Mermaid diagram in the README shows the pipeline flow, which this PR doesn't change. |
| Changelog gate | ✅ | The entry is updated with the new changes. check_changelog OK locally. |
| Observability | ✅ | No new stage. ⏭️ The optional reasoningEffort in *.llm_request meta is not added: it was optional and outside scope. |
Suggested before merge
Check the Groq facts (Warning 3).✔️ Cross-checked (see W3).Set✔️ Done (W1).max_tokenson the 3 uncapped stages.Update the ADR-0005/0017 statuses and✔️ Done.AGENTS.md:25.Resolve the merge conflict with✔️ Done inmain(feat(review): ground PR review verdict in executed tool evidence (ADR-0024) #160).41fa5f8(Update 2).- After merge: push to an open PR and confirm the
reviewjob gets past the LLM call (no 404/413/400). Still to do; it can only be checked after merge (ADR-0023).
Generated by Claude Code
…T override Addresses review on #161: - Explicit <stage>_max_tokens for validation (1024), generation (4096), review (1024). - autofix_max_input_tokens 3400 -> 3000: auto-fix-system.md measures ~890 tokens, not ~460, so the previous budget totaled ~8,385 > 8K TPM. Budget test now measures the prompt with estimateTokens() instead of a hard-coded constant. - GROQ_REASONING_EFFORT repo variable overrides every stage; `off` drops the parameter for non-reasoning GROQ_MODEL overrides. Wired into all 4 workflows. - ADR-0005 superseded, ADR-0017 amended; ADR-0024, AGENTS (stale 32K context line), code-generation (per-stage keys table, env matrix), runbook (400 row). - Tests: validate_issue/generate_issue_change wiring E2E, global fallback key, env override/off/invalid, Anthropic ignores reasoningEffort, explicit max_tokens on every stage. Retired-model test no longer pins the exact ID. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQ2EvwfnVxY9PXffXtc3S4
…to 0025 #160 took ADR-0024 (tool evidence for PR review). The Groq model ADR becomes ADR-0025: file renamed, index, ADR-0005/0017 status links and every Groq reference updated. pr_review.test.mjs conflict resolved by keeping both sides (44 tests = 36 base + 1 here + 7 from #160). ADR-0025 and the runbook 413 row now account for the tool-evidence block (up to 2,000 chars per failing check) added to the review prompt by #160, which reduces the 8K TPM headroom of the review stage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013b1tRLHJqkbkfEZKwonphu
🎯 Goal
Restores the Groq path. Both default models in
config/models.yamlhave been retired by Groq, so every Groq call currently fails with404 model_not_found(seen on thereviewjob of #159).qwen/qwen3-32bllama-3.3-70b-versatile📦 Changes
openai/gpt-oss-120b, Groq's recommended replacement.GROQ_MODELstill overrides every stage.max_tokensand TPM.<stage>_reasoning_effort(low|medium|high), validated byloadLLMConfig().loadConfigFromEnv(), which previously dropped unlisted fields.groq_clientsends it only when it is set, because non-reasoning models reject the parameter.low.GROQ_REASONING_EFFORToverrides every stage;offstops sending the parameter (for a non-reasoningGROQ_MODEL).validation_max_tokens: 1024,generation_max_tokens: 4096,review_max_tokens: 1024.autofix_max_input_tokensgoes from 7,400 to 3,000, so that system (~890, now measured withestimateTokens()) + input + 4,096 output ≈ 7,986 fits the 8K free-tier TPM.auto_fix_prnow knows the model's 131,072-token context window.docs/code-generation.mdand the runbook (rows for 404model_not_found, 400reasoning_effort, 413).🧪 Validation
main. Lint, the c8 coverage gate onconfig.mjsand the changelog check pass.max_tokens.reasoningEffortis loaded and validated, including the absent-key, global-fallback and invalid-value branches, and theGROQ_REASONING_EFFORToverride/off/empty/invalid cases.loadConfigFromEnvforwards it.callGroqsends or omitsreasoning_effort;callAnthropicignores it.model: openai/gpt-oss-120b,max_tokensandreasoning_effort: low.autofix_max_input_tokens.review(~6.5K prompt tokens + 1,024 output) andgeneration(no input cap) can hit 413/429 on large PRs or issues.main's config, so thereviewcheck on this PR will still use the retired model and fail. The new config only takes effect after merge.🔁 Checklist
🧪 How to test
npm testreviewjob should call Groq withopenai/gpt-oss-120band get past the LLM call.🤖 Generated with Claude Code
https://claude.ai/code/session_013b1tRLHJqkbkfEZKwonphu