Skip to content

[agentic-token-optimizer] Optimize Avenger workflow: fix model mismatch driving high AIC #50312

Description

@github-actions

Target workflow

Avenger (.github/workflows/avenger.md) — hourly CI-fixer, top workflow by total AIC in the last 7 days (3 runs, no ## agent: blocks, not optimized in the last 14 days).

Analysis period

7-day window, all 3 runs analyzed (100% of available runs) using all-runs.json plus local per-run artifacts under .github/aw/logs/run-<id>/agent_usage.json and run_summary.json.

Cost profile

Metric Value
Total AIC (3 runs) 322.557
Avg AIC / run 107.519
Total raw tokens 34,332
Avg action minutes / run 14.7
Turns reported 0 (not tracked by engine)
Errors / warnings 0

Root cause: model mismatch

The workflow frontmatter pins:

model: claude-haiku-4.5
max-turns: 50

But every one of the 3 audited runs actually executed on claude-opus-4-8 per agent_usage.json ("primary_model":"claude-opus-4-8"), not haiku:

Run AIC primary_model cache_read_tokens
§30922901541 94.872 claude-opus-4-8 925,147
§30917409088 124.255 claude-opus-4-8 1,352,166
§30912183153 103.430 claude-opus-4-8 999,167

Per models.json cost tables, opus-4-8 is priced at input $5e-6/output $2.5e-5/cache-read $5e-7, vs. haiku-4-5 at input $1e-6/output $5e-6/cache-read $1e-7 — roughly 5× more expensive per token across the board. The raw token counts (10–13k) are modest for the task, but AIC is driven almost entirely by very large cache-read volumes (0.9M–1.35M tokens) billed at opus rates.

Ranked recommendations

1. Confirm/enforce the configured model is actually used — expected savings: ~70–80% of AIC/run (~75–100 AIC/run)

  • Action: Verify why model: claude-haiku-4.5 in frontmatter is not honored at runtime (engine/version mismatch, or model override elsewhere in the awf/engine layer). File a compiler/engine bug if the declared model isn't being passed through, or pin the model more explicitly (e.g., exact model ID claude-haiku-4-5 instead of an alias) and re-verify via aw_info.json.model matching agent_usage.json.primary_model after the next few runs.
  • Evidence: all 3 runs show aw_info.json.model = "claude-haiku-4.5" but agent_usage.json.primary_model = "claude-opus-4-8" — a total mismatch that likely accounts for the majority of the AIC on this workflow.

2. Reduce cache-read volume via a lighter Step 0 CI check — estimated savings: 10–15% AIC/run

  • Action: The check_ci_status job already fetches CI state via gh run list in a plain GitHub Actions step (not the agent). Step 0 in the agent prompt re-verifies live CI status with another gh run list call inside the agent loop. Since the composite CI status is already passed in via ${{ needs.check_ci_status.outputs.ci_status }}, consider skipping the redundant live re-check unless the passed-in status is stale/ambiguous, cutting one round-trip of context re-loading.
  • Evidence: prompt.txt Step 0 duplicates the check_ci_status job's gh run list --workflow=ci.yml --branch=main --limit=2 call; this is a repeated tool call whose result was already available from the job output.

3. Tighten max-turns ceiling for a narrow, single-purpose repair loop — estimated savings: 3–5% AIC/run on outlier runs

  • Action: max-turns: 50 is generous for a workflow whose own guidance states "Hard limit is 25 turns." Lower max-turns to ~25–30 to match the documented budget and cap outlier runs (the highest-cost run, §30917409088, also had the highest cache-read volume).
  • Evidence: workflow's own "Execution Guidelines" section says "Token Budget Awareness: Hard limit is 25 turns" while frontmatter allows up to 50.

Caveats

  • Only 3 runs were available in the 7-day window — a small sample. Recommendation rejig docs #1 should be re-validated against the next 5–10 runs once (if) the model configuration is corrected, to confirm the AIC drop materializes.
  • Turns and ToolCalls were not populated in the metrics for this engine/version, so tool-usage-pattern analysis (Phase 2 tool table) could not be completed with call-level granularity; only prompt-level duplication (Step 0 vs. check_ci_status job) could be identified from prompt.txt.
  • No structural setup-prefix or inline sub-agent optimization is recommended: the workflow has a single linear repair sequence (merge → recompile → fmt → wasm-golden → lint → test) where each step depends on the prior step's working tree state, so sections are not independent enough for parallel extraction, and the workflow already has no ## agent: blocks but the task is too sequential/stateful for the 3-section-independence bar in Phase 4.

References: §30922901541, §30917409088, §30912183153

Generated by Agentic Workflow AIC Usage Optimizer · auto · 74.9 AIC · ⊞ 10.5K · ◷

  • expires on Aug 11, 2026, 7:46 AM UTC-08:00

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions