Upgrade the self-hosted PR-review engine's deep tier to Claude Opus 4.8 (shipped 2026-05-28) for its honesty/quality gains (4x less likely to miss flawed code) and its model-internal Dynamic Workflows on large diffs. The deep tier (ENGINE_DEEP_MODEL) and its primary fallback chain are already pinned to claude-opus-4-8 in scripts/engine.sh, so this initiative is about (a) validating that the live upgrade is non-regressive against the frozen held-out eval set before we declare it done, (b) finishing the half-done reference rollout (stale claude-opus-4-7 entries still linger in fallback chains, scripts/aw.sh, and docs/dev-lead/spec.md), and (c) keeping model-pricing.tsv correct for the new model id.
Scope decision — NO Express/Fast tier. The idea proposed a Fast Mode "express" tier at the new $10/$50 per-MTok rate. The repo maintainer rejected it (discussion #793 comment): "do not want to enable the fast tier as we do not have any need for an Express PR review capability. The bottleneck in PR reviews is not the speed of the model." This initiative therefore touches only standard Opus 4.8 routing and pricing; it adds no express tier and no Fast Mode pricing.
Initiative success metric (how we know it worked): the deep-review held-out eval suite (evals/deep-review/holdout/cases.jsonl, scored by scripts/evals/run-eval.sh in llm-judge mode) run under claude-opus-4-8 yields an aggregate score >= the recorded claude-opus-4-7 baseline on the identical frozen set (strict non-regression). The baseline score is recorded as an immutable artifact before the candidate run. A regression is a STOP signal: the deep tier must be reverted to claude-opus-4-7 rather than the rollout being finished.
Cost cap / budget bound. Validation is bounded to the 5-case deep-review holdout set, run via run_agentic (Opus deep tier) for at most 2 models (4.7 baseline + 4.8 candidate) x at most 2 repeats = <= 20 Opus deep invocations total, with TOKEN_LOG_FILE capture so the dollar cost is reported per AGENTS.md cost reporting. Steady-state review cost is unchanged: claude-opus-4-8 is priced by the same claude-opus-4-* row ($5 / $0.50 / $6.25 / $25 per MTok, effective 2025-11-01) that already priced 4.7 — the upgrade is cost-neutral per review. The rejected Fast Mode rate ($10/$50) is explicitly NOT introduced.
Untracked prerequisites
Planned from idea discussion #793 by the BMAD Scrum Master initiative-planner. Inert until a maintainer adds initiative:auto.
Upgrade the self-hosted PR-review engine's deep tier to Claude Opus 4.8 (shipped 2026-05-28) for its honesty/quality gains (4x less likely to miss flawed code) and its model-internal Dynamic Workflows on large diffs. The deep tier (
ENGINE_DEEP_MODEL) and its primary fallback chain are already pinned toclaude-opus-4-8inscripts/engine.sh, so this initiative is about (a) validating that the live upgrade is non-regressive against the frozen held-out eval set before we declare it done, (b) finishing the half-done reference rollout (staleclaude-opus-4-7entries still linger in fallback chains,scripts/aw.sh, anddocs/dev-lead/spec.md), and (c) keepingmodel-pricing.tsvcorrect for the new model id.Scope decision — NO Express/Fast tier. The idea proposed a Fast Mode "express" tier at the new $10/$50 per-MTok rate. The repo maintainer rejected it (discussion #793 comment): "do not want to enable the fast tier as we do not have any need for an Express PR review capability. The bottleneck in PR reviews is not the speed of the model." This initiative therefore touches only standard Opus 4.8 routing and pricing; it adds no express tier and no Fast Mode pricing.
Initiative success metric (how we know it worked): the
deep-reviewheld-out eval suite (evals/deep-review/holdout/cases.jsonl, scored byscripts/evals/run-eval.shin llm-judge mode) run underclaude-opus-4-8yields an aggregate score >= the recordedclaude-opus-4-7baseline on the identical frozen set (strict non-regression). The baseline score is recorded as an immutable artifact before the candidate run. A regression is a STOP signal: the deep tier must be reverted toclaude-opus-4-7rather than the rollout being finished.Cost cap / budget bound. Validation is bounded to the 5-case
deep-reviewholdout set, run viarun_agentic(Opus deep tier) for at most 2 models (4.7 baseline + 4.8 candidate) x at most 2 repeats = <= 20 Opus deep invocations total, withTOKEN_LOG_FILEcapture so the dollar cost is reported per AGENTS.md cost reporting. Steady-state review cost is unchanged:claude-opus-4-8is priced by the sameclaude-opus-4-*row ($5 / $0.50 / $6.25 / $25 per MTok, effective 2025-11-01) that already priced 4.7 — the upgrade is cost-neutral per review. The rejected Fast Mode rate ($10/$50) is explicitly NOT introduced.Untracked prerequisites
claude-opus-4-8) must be reachable from the review/CI runtime via theclaudeCLI / Anthropic API with the available credentials — an external availability dependency, not a repo issue. If the model id is not yet served to this account, the eval-validation story cannot run.Planned from idea discussion #793 by the BMAD Scrum Master initiative-planner. Inert until a maintainer adds
initiative:auto.