Skip to content

[FEATURE] Cap the PR review input tokens (review_max_input_tokens) to stay under Groq's 8K TPM #163

Description

@koydas

🎯 Goal

Bound the size of the review LLM request so it always fits Groq's per-request TPM limit (8,000 on the free tier for openai/gpt-oss-120b, ADR-0025), the same way autofix_max_input_tokens already does for auto-fix (ADR-0017).

📍 Context

🚀 Description

Today the only cap on the review prompt is the diff's 12,000 chars (filterDiff). The PR body, dependency manifests, change classification, automation-gate context and the tool evidence from ADR-0024 are uncapped.

Measured on #161: input_tokens_est: 6743, plus review_max_tokens: 1024, totals ~7,767. That leaves about 230 tokens of margin under 8K. A PR with failing checks adds up to 2,000 chars of output tail per failing check (~500 tokens each), so it is the most likely case to get a 413 Request too large. After a 413 the review fails and no verdict is posted.

Proposal:

  • New optional key review_max_input_tokens, validated by loadLLMConfig() like autofix_max_input_tokens. Default it so that system + input + review_max_tokens ≤ 8,000.

  • pr_review.mjs assembles the user prompt within that budget, trimming in priority order:

    1. manifests
    2. PR body
    3. diff

    Tool evidence and the verdict instructions are never trimmed, since ADR-0024 relies on them.

  • When anything is trimmed, the prompt states it, as the existing diffTruncated warning does.

  • Log section token estimates in review.llm_request meta, like auto-fix's token_estimate.

🧩 Scope

  • In: review stage budget, the config key, docs (docs/code-generation.md per-stage keys table, runbook 413 row), and a config test asserting the sum ≤ 8,000 using the prompt measured with estimateTokens().
  • Out: generation stage input cap (separate issue if needed), tokenizer-accurate counting, Developer-plan tuning.

🧪 Acceptance criteria

  • review_max_input_tokens is loaded and validated: absent, valid, and invalid (non-positive, non-integer) branches are tested.
  • With a huge PR body, a large diff and 2 failing checks, the request built by pr_review.mjs stays ≤ review_max_input_tokens. Tool evidence is intact and the trim notice is present.
  • The test fails if the system prompt + review_max_input_tokens + review_max_tokens > 8,000.
  • Removing the key restores current behavior (no cap).
  • Docs and CHANGELOG are updated. Amend ADR-0025 or add a new ADR.

⚙️ Constraints

  • Must not drop or truncate the tool-evidence block or the verdict format instructions (ADR-0024).
  • No new dependencies; chars/4 estimation (estimateTokens) is acceptable.
  • Edit-guardrails: targeted edit in pr_review.mjs (≤ 30% of lines).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ready-for-devIssue validated and ready for automated implementation

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions