🎯 Goal
Bound the size of the review LLM request so it always fits Groq's per-request TPM limit (8,000 on the free tier for openai/gpt-oss-120b, ADR-0025), the same way autofix_max_input_tokens already does for auto-fix (ADR-0017).
📍 Context
🚀 Description
Today the only cap on the review prompt is the diff's 12,000 chars (filterDiff). The PR body, dependency manifests, change classification, automation-gate context and the tool evidence from ADR-0024 are uncapped.
Measured on #161: input_tokens_est: 6743, plus review_max_tokens: 1024, totals ~7,767. That leaves about 230 tokens of margin under 8K. A PR with failing checks adds up to 2,000 chars of output tail per failing check (~500 tokens each), so it is the most likely case to get a 413 Request too large. After a 413 the review fails and no verdict is posted.
Proposal:
-
New optional key review_max_input_tokens, validated by loadLLMConfig() like autofix_max_input_tokens. Default it so that system + input + review_max_tokens ≤ 8,000.
-
pr_review.mjs assembles the user prompt within that budget, trimming in priority order:
- manifests
- PR body
- diff
Tool evidence and the verdict instructions are never trimmed, since ADR-0024 relies on them.
-
When anything is trimmed, the prompt states it, as the existing diffTruncated warning does.
-
Log section token estimates in review.llm_request meta, like auto-fix's token_estimate.
🧩 Scope
- In:
review stage budget, the config key, docs (docs/code-generation.md per-stage keys table, runbook 413 row), and a config test asserting the sum ≤ 8,000 using the prompt measured with estimateTokens().
- Out:
generation stage input cap (separate issue if needed), tokenizer-accurate counting, Developer-plan tuning.
🧪 Acceptance criteria
⚙️ Constraints
- Must not drop or truncate the tool-evidence block or the verdict format instructions (ADR-0024).
- No new dependencies; chars/4 estimation (
estimateTokens) is acceptable.
- Edit-guardrails: targeted edit in
pr_review.mjs (≤ 30% of lines).
🎯 Goal
Bound the size of the
reviewLLM request so it always fits Groq's per-request TPM limit (8,000 on the free tier foropenai/gpt-oss-120b, ADR-0025), the same wayautofix_max_input_tokensalready does for auto-fix (ADR-0017).📍 Context
scripts/pr_review.mjs,config/models.yaml,scripts/lib/config.mjs🚀 Description
Today the only cap on the review prompt is the diff's 12,000 chars (
filterDiff). The PR body, dependency manifests, change classification, automation-gate context and the tool evidence from ADR-0024 are uncapped.Measured on #161:
input_tokens_est: 6743, plusreview_max_tokens: 1024, totals ~7,767. That leaves about 230 tokens of margin under 8K. A PR with failing checks adds up to 2,000 chars of output tail per failing check (~500 tokens each), so it is the most likely case to get a 413 Request too large. After a 413 the review fails and no verdict is posted.Proposal:
New optional key
review_max_input_tokens, validated byloadLLMConfig()likeautofix_max_input_tokens. Default it so that system + input +review_max_tokens≤ 8,000.pr_review.mjsassembles the user prompt within that budget, trimming in priority order:Tool evidence and the verdict instructions are never trimmed, since ADR-0024 relies on them.
When anything is trimmed, the prompt states it, as the existing
diffTruncatedwarning does.Log section token estimates in
review.llm_requestmeta, like auto-fix'stoken_estimate.🧩 Scope
reviewstage budget, the config key, docs (docs/code-generation.mdper-stage keys table, runbook 413 row), and a config test asserting the sum ≤ 8,000 using the prompt measured withestimateTokens().generationstage input cap (separate issue if needed), tokenizer-accurate counting, Developer-plan tuning.🧪 Acceptance criteria
review_max_input_tokensis loaded and validated: absent, valid, and invalid (non-positive, non-integer) branches are tested.pr_review.mjsstays ≤review_max_input_tokens. Tool evidence is intact and the trim notice is present.review_max_input_tokens+review_max_tokens> 8,000.⚙️ Constraints
estimateTokens) is acceptable.pr_review.mjs(≤ 30% of lines).