Skip to content

Opus 4.8 engine upgrade: validate the deep-review tier, finish the reference rollout, and keep cost data accurate #835

Description

@github-actions

Upgrade the self-hosted PR-review engine's deep tier to Claude Opus 4.8 (shipped 2026-05-28) for its honesty/quality gains (4x less likely to miss flawed code) and its model-internal Dynamic Workflows on large diffs. The deep tier (ENGINE_DEEP_MODEL) and its primary fallback chain are already pinned to claude-opus-4-8 in scripts/engine.sh, so this initiative is about (a) validating that the live upgrade is non-regressive against the frozen held-out eval set before we declare it done, (b) finishing the half-done reference rollout (stale claude-opus-4-7 entries still linger in fallback chains, scripts/aw.sh, and docs/dev-lead/spec.md), and (c) keeping model-pricing.tsv correct for the new model id.

Scope decision — NO Express/Fast tier. The idea proposed a Fast Mode "express" tier at the new $10/$50 per-MTok rate. The repo maintainer rejected it (discussion #793 comment): "do not want to enable the fast tier as we do not have any need for an Express PR review capability. The bottleneck in PR reviews is not the speed of the model." This initiative therefore touches only standard Opus 4.8 routing and pricing; it adds no express tier and no Fast Mode pricing.

Initiative success metric (how we know it worked): the deep-review held-out eval suite (evals/deep-review/holdout/cases.jsonl, scored by scripts/evals/run-eval.sh in llm-judge mode) run under claude-opus-4-8 yields an aggregate score >= the recorded claude-opus-4-7 baseline on the identical frozen set (strict non-regression). The baseline score is recorded as an immutable artifact before the candidate run. A regression is a STOP signal: the deep tier must be reverted to claude-opus-4-7 rather than the rollout being finished.

Cost cap / budget bound. Validation is bounded to the 5-case deep-review holdout set, run via run_agentic (Opus deep tier) for at most 2 models (4.7 baseline + 4.8 candidate) x at most 2 repeats = <= 20 Opus deep invocations total, with TOKEN_LOG_FILE capture so the dollar cost is reported per AGENTS.md cost reporting. Steady-state review cost is unchanged: claude-opus-4-8 is priced by the same claude-opus-4-* row ($5 / $0.50 / $6.25 / $25 per MTok, effective 2025-11-01) that already priced 4.7 — the upgrade is cost-neutral per review. The rejected Fast Mode rate ($10/$50) is explicitly NOT introduced.

Untracked prerequisites

  • Claude Opus 4.8 (claude-opus-4-8) must be reachable from the review/CI runtime via the claude CLI / Anthropic API with the available credentials — an external availability dependency, not a repo issue. If the model id is not yet served to this account, the eval-validation story cannot run.
  • discussion 💡 Opus 4.8 Engine Upgrade: Dynamic Workflows + 4x Honesty for Review Quality #793 comment from the maintainer rejecting the Fast Mode / Express review tier is a recorded scope decision this plan honors (no express tier, no Fast Mode pricing).

Planned from idea discussion #793 by the BMAD Scrum Master initiative-planner. Inert until a maintainer adds initiative:auto.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    initiativeEpic / initiative tracking issueinitiative:autoDriver may auto-release this epic's ready sub-issues to dev-lead

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions