Problem
No published measurement covers Claude 5.5 models on this repo's subagent roles: codebase retrieval, code review and verification, and bounded mechanical edits. Routing defaults in plugins/multi-agent/reference/defaults.yaml therefore rest on judgment. Haiku 5.5 may take no role until that role passes this eval.
Proposal
After the Haiku 5.5 documentation PRs land, build a paired eval comparing these setups on the same tasks from this repo:
- Haiku 5.5 at
high and xhigh
- Haiku 5.5 with an Opus 5.5 advisor, using Anthropic's Haiku coding prompt
- Sonnet 5.5 at
low, medium and high
- Opus 5.5 at
low and medium
Measure:
- pass rate, or verified findings
- tokens per completed task (input, output and thinking)
- advisor call rate
- wall time
State no prices. Use a judge from outside the candidate set, or Opus with its self-preference noted.
Acceptance
Problem
No published measurement covers Claude 5.5 models on this repo's subagent roles: codebase retrieval, code review and verification, and bounded mechanical edits. Routing defaults in
plugins/multi-agent/reference/defaults.yamltherefore rest on judgment. Haiku 5.5 may take no role until that role passes this eval.Proposal
After the Haiku 5.5 documentation PRs land, build a paired eval comparing these setups on the same tasks from this repo:
highandxhighlow,mediumandhighlowandmediumMeasure:
State no prices. Use a judge from outside the candidate set, or Opus with its self-preference noted.
Acceptance