Daily Spending Forecast — github/gh-aw
Forecast date: 2026-08-03 · History window: 30 days (2026-07-04 → 2026-08-03) · Period basis: month
Run reference: §30807125030
Executive Summary
- Total observed AIC (30d, all sampled runs): 46,444.36
- Active (AIC-tracked) workflows: 23 of 49 total workflows discovered
- Weekly forecast: P10 4,216.50 · P50 9,602.17 · P90 17,376.99
- Monthly forecast: P10 27,359.71 · P50 42,724.51 · P90 61,634.04
- Top spenders by observed AIC: Go Logger Enhancement (14,197.7), Agentic Workflow Audit Agent (7,211.3), Smoke Copilot (5,997.6) — together ~59% of observed spend.
- 26 of 49 workflows show
sampled_runs = 0; these are non-agentic/CI workflows (CodeQL, Dependabot, Format/Lint/Build, Mergefest, test harnesses, etc.) that either had zero AIC-tracked runs in the window or produced AIC=0 (treated as missing data), not a data-collection failure — confirmed via debug log (see Data Quality section).
Daily Spending Trend (observed AIC, all workflows, 2026-07-04 → 2026-08-03)
07-04 |################################# 1480
07-05 |######################### 1144
07-06 |############################################### 2120
07-07 |############################## 1347
07-08 |################################################ 2174
07-09 |###################################### 1690
07-10 |################################### 1590
07-11 |#################################### 1636
07-12 |#################################### 1631
07-13 |################################################ 2138
07-14 |############################# 1296
07-15 |########################################## 1885
07-16 |##################################### 1672
07-17 |################################## 1530
07-18 |################################## 1529
07-19 |############################################# 2026
07-20 |######################################## 1792
07-21 |################################################## 2223
07-22 |################################################# 2218
07-23 |################################# 1483
07-24 |##################################### 1684
07-25 |################################# 1492
07-26 |### 168
07-27 |############ 565
07-28 |###################### 989
07-29 |################################# 1481
07-30 |#################### 930
07-31 |################### 857
08-01 |################################### 1598
08-02 |########################### 1210
08-03 |################### 865
Note the sharp dip 07-26 → 07-31 (avg ~752/day vs ~1,760/day baseline). No corresponding
error entries were found in the diagnostics logs for that window; this is likely reduced
weekday/weekend or PR-driven trigger volume rather than a collection gap — flagged below.
Forecast Confidence Chart (aggregate P10/P50/P90 AIC)
Weekly P10 |## 4216
Weekly P50 |###### 9602
Weekly P90 |########### 17377
Monthly P10 |################# 27360
Monthly P50 |########################### 42725
Monthly P90 |######################################## 61634
Workflow Forecast Table (active / AIC-tracked workflows, sorted by observed AIC)
| Workflow |
Samples |
Observed AIC |
P50/run |
P95/run |
Weekly P50 |
Monthly P50 |
Success |
Monthly CI (P10–P90) |
| Go Logger Enhancement |
28 |
14,197.7 |
504.48 |
757.24 |
3,123.1 |
13,734.9 |
96% |
9,119 – 19,004 |
| Agentic Workflow Audit Agent |
27 |
7,211.3 |
266.25 |
400.11 |
1,526.9 |
6,682.1 |
93% |
— |
| Smoke Copilot |
70 |
5,997.6 |
86.71 |
129.56 |
948.1 |
4,097.6 |
69% |
— |
| CLI Version Checker |
28 |
3,655.3 |
113.87 |
213.36 |
770.4 |
3,389.6 |
93% |
— |
| Copilot Agent PR Analysis |
27 |
2,710.2 |
91.82 |
170.82 |
624.8 |
2,713.3 |
100% |
— |
| Lockfile Statistics Analysis Agent |
27 |
1,806.3 |
59.30 |
131.61 |
411.3 |
1,818.8 |
100% |
— |
| Terminal Stylist |
31 |
1,660.1 |
46.81 |
97.12 |
376.5 |
1,660.1 |
100% |
— |
| Smoke Claude |
17 |
1,297.5 |
75.89 |
105.57 |
270.7 |
1,225.5 |
94% |
— |
| Tidy |
46 |
1,292.6 |
21.66 |
70.72 |
272.7 |
1,210.4 |
93% |
— |
| GitHub MCP Remote Server Tools Report Generator |
4 |
1,192.1 |
244.29 |
515.92 |
160.6 |
837.1 |
75% |
sparse sample |
| Scout |
8 |
1,125.7 |
86.95 |
328.05 |
181.0 |
1,100.5 |
100% |
— |
| Weekly Workflow Analysis |
4 |
876.0 |
215.51 |
224.17 |
215.5 |
875.2 |
100% |
sparse sample |
| Daily Documentation Updater |
29 |
717.0 |
22.29 |
42.93 |
163.7 |
724.2 |
100% |
— |
| Daily News |
20 |
705.1 |
31.33 |
52.82 |
160.0 |
706.7 |
100% |
— |
| Dev |
31 |
416.3 |
14.12 |
22.80 |
94.1 |
417.7 |
100% |
— |
| Documentation Unbloat |
28 |
349.8 |
11.80 |
— |
76.5 |
337.3 |
96% |
— |
| Duplicate Code Detector |
30 |
275.0 |
7.05 |
— |
61.1 |
272.3 |
100% |
— |
| Weekly Issue Summary |
4 |
205.2 |
37.94 |
— |
37.9 |
205.0 |
100% |
sparse sample |
| Smoke Codex |
17 |
197.9 |
13.66 |
— |
35.4 |
162.0 |
82% |
— |
| Plan Command |
6 |
187.1 |
20.87 |
— |
28.2 |
185.3 |
100% |
sparse sample |
| Artifacts Usage Report |
5 |
177.1 |
38.65 |
— |
38.7 |
177.1 |
100% |
sparse sample |
| MCP Inspector Agent |
2 |
115.2 |
42.89 |
— |
0.0 |
115.2 |
100% |
sparse sample |
| Repository Tree Map Generator |
3 |
76.4 |
25.21 |
— |
25.1 |
76.4 |
100% |
sparse sample |
Bold sample counts (≤5 runs) indicate low statistical confidence; treat their weekly/monthly
projections as directional only.
Data quality and accuracy notes (31 flags detected)
Sparse samples (≤5 runs in 30 days) — low confidence projections:
MCP Inspector Agent (2 runs), Repository Tree Map Generator (3 runs),
Weekly Workflow Analysis, GitHub MCP Remote Server Tools Report Generator,
Weekly Issue Summary (4 runs each), Plan Command, Artifacts Usage Report (5–6 runs).
These are expected given weekly/on-demand trigger cadence rather than a data gap, but
their Monte Carlo intervals are wide relative to their mean and should not anchor a
hard budget commitment. MCP Inspector Agent's weekly P50 rounds to 0.0 because its
2 samples fall outside the trailing 7-day window used for that projection — this is a
windowing artifact, not zero spend.
Zero-sampled-run workflows (26 of 49): Copilot Setup Steps, Video Analysis Agent,
Mergefest, Notion Issue Summary, Dev Hawk, Poem Bot - A Creative Agentic Workflow,
Q, Rebuild the documentation after making changes, Go Pattern Detector,
Resource Summarizer Agent, Dependabot Updates, Sentry Issue Analyzer,
Doc Build - Deploy, Format, Lint, Build and Commit, Test Claude, Smoke OpenCode,
CodeQL, Test, Commit Changes Analyzer, Test Copilot CLI Engine,
Test Copilot GitHub Integration, CI Failure Doctor, .github/workflows/test-proxy,
CI, Basic Research Agent, copilot only.
Follow-up: inspected forecast.stderr.log debug output for CI Failure Doctor (a
representative sample) and confirmed the tool explicitly detected and skipped 10 runs
with AIC=0.000 treated as missing data, then logged No non-zero AIC run samples found ... in last 30 days. This confirms the zero counts are a correct reflection of
non-agentic/CI-only workflows (build, lint, CodeQL, Dependabot, smoke/test harnesses)
rather than an artifact-download or collection failure. The command exited 0 with no
error entries in the transcript. Resolved — no rerun needed.
Daily spending dip (2026-07-26 to 2026-07-31, ~57% below 30-day average): No error or
warning log entries correspond to this window. Likely explanation is reduced PR/issue
volume over that period (fewer trigger events for on-demand/PR-driven workflows) rather
than a sampling defect, but this is not independently confirmed and should be treated as
an open question if forecast accuracy for early-August budgeting is critical.
Wide confidence intervals: Workflows with sparse samples (flagged above) naturally
produce wide P10–P90 spreads relative to their P50 (a Monte Carlo artifact of small
sample size, not a data error). All are marked is_reliable: true in the raw JSON,
but "reliable" here means the Monte Carlo procedure ran successfully, not that the
few underlying samples are representative — use these figures as budget minimums/
maximums, not as SLA-style guarantees.
No AIC values of exactly 0 appear in the retained run_samples arrays for the 23
active workflows — the pipeline appears to have already filtered zero-AIC runs before
populating the JSON (consistent with the treated as missing data filtering logic seen
in the debug log for CI Failure Doctor).
History window and date consistency: All 23 active workflows report history_days: 30 and sample dates falling within 2026-07-04–2026-08-03, consistent with the stated
window; no stale or out-of-range samples were found.
Assumptions
- Forecast basis:
month period, 30-day trailing history, as generated by gh aw forecast
at 2026-08-03T10:54:31Z with exit code 0 (no rerun was necessary).
- "Observed AIC" = sum of actual
run_samples[].aic values recorded in the last 30 days;
"projected AIC" = Monte-Carlo-derived P10/P50/P90 extrapolations for the next
week/month, not historical totals.
- Zero-sampled-run workflows are treated as having no measurable agentic cost in this
window (build/lint/CI/test/Dependabot-type jobs) and are excluded from the observed and
projected totals.
- Sparse-sample workflows (≤5 runs) are included in totals but flagged as low-confidence;
their contribution to the aggregate P10/P50/P90 is small (<3% of the monthly P50 total).
Next Actions
- Treat the monthly P50 ≈ 42,725 AIC as the primary budget planning figure; use
P90 ≈ 61,634 as a conservative ceiling for approvals.
- Investigate the 2026-07-26–07-31 spending dip if tighter short-term forecasting is
needed (check for reduced PR volume or a scheduling gap in that window).
- Continue monitoring
Go Logger Enhancement, Agentic Workflow Audit Agent, and
Smoke Copilot — they account for ~59% of observed spend and dominate forecast
variance.
- For sparse-sample workflows, consider
gh aw forecast --eval backtesting once more
history accumulates (target ≥10 samples) before relying on their individual P10/P90
ranges for budgeting decisions.
Generated by 📈 Daily Spending Forecast · auto · 44.1 AIC · ⌖ 3.73 AIC · ⊞ 9.1K · ◷
Daily Spending Forecast — github/gh-aw
Forecast date: 2026-08-03 · History window: 30 days (2026-07-04 → 2026-08-03) · Period basis: month
Run reference: §30807125030
Executive Summary
sampled_runs = 0; these are non-agentic/CI workflows (CodeQL, Dependabot, Format/Lint/Build, Mergefest, test harnesses, etc.) that either had zero AIC-tracked runs in the window or produced AIC=0 (treated as missing data), not a data-collection failure — confirmed via debug log (see Data Quality section).Daily Spending Trend (observed AIC, all workflows, 2026-07-04 → 2026-08-03)
Note the sharp dip 07-26 → 07-31 (avg ~752/day vs ~1,760/day baseline). No corresponding
error entries were found in the diagnostics logs for that window; this is likely reduced
weekday/weekend or PR-driven trigger volume rather than a collection gap — flagged below.
Forecast Confidence Chart (aggregate P10/P50/P90 AIC)
Workflow Forecast Table (active / AIC-tracked workflows, sorted by observed AIC)
Bold sample counts (≤5 runs) indicate low statistical confidence; treat their weekly/monthly
projections as directional only.
Data quality and accuracy notes (31 flags detected)
Sparse samples (≤5 runs in 30 days) — low confidence projections:
MCP Inspector Agent(2 runs),Repository Tree Map Generator(3 runs),Weekly Workflow Analysis,GitHub MCP Remote Server Tools Report Generator,Weekly Issue Summary(4 runs each),Plan Command,Artifacts Usage Report(5–6 runs).These are expected given weekly/on-demand trigger cadence rather than a data gap, but
their Monte Carlo intervals are wide relative to their mean and should not anchor a
hard budget commitment.
MCP Inspector Agent's weekly P50 rounds to 0.0 because its2 samples fall outside the trailing 7-day window used for that projection — this is a
windowing artifact, not zero spend.
Zero-sampled-run workflows (26 of 49):
Copilot Setup Steps,Video Analysis Agent,Mergefest,Notion Issue Summary,Dev Hawk,Poem Bot - A Creative Agentic Workflow,Q,Rebuild the documentation after making changes,Go Pattern Detector,Resource Summarizer Agent,Dependabot Updates,Sentry Issue Analyzer,Doc Build - Deploy,Format, Lint, Build and Commit,Test Claude,Smoke OpenCode,CodeQL,Test,Commit Changes Analyzer,Test Copilot CLI Engine,Test Copilot GitHub Integration,CI Failure Doctor,.github/workflows/test-proxy,CI,Basic Research Agent,copilot only.Follow-up: inspected
forecast.stderr.logdebug output forCI Failure Doctor(arepresentative sample) and confirmed the tool explicitly detected and skipped 10 runs
with
AIC=0.000 treated as missing data, then loggedNo non-zero AIC run samples found ... in last 30 days. This confirms the zero counts are a correct reflection ofnon-agentic/CI-only workflows (build, lint, CodeQL, Dependabot, smoke/test harnesses)
rather than an artifact-download or collection failure. The command exited 0 with no
error entries in the transcript. Resolved — no rerun needed.
Daily spending dip (2026-07-26 to 2026-07-31, ~57% below 30-day average): No error or
warning log entries correspond to this window. Likely explanation is reduced PR/issue
volume over that period (fewer trigger events for on-demand/PR-driven workflows) rather
than a sampling defect, but this is not independently confirmed and should be treated as
an open question if forecast accuracy for early-August budgeting is critical.
Wide confidence intervals: Workflows with sparse samples (flagged above) naturally
produce wide P10–P90 spreads relative to their P50 (a Monte Carlo artifact of small
sample size, not a data error). All are marked
is_reliable: truein the raw JSON,but "reliable" here means the Monte Carlo procedure ran successfully, not that the
few underlying samples are representative — use these figures as budget minimums/
maximums, not as SLA-style guarantees.
No AIC values of exactly 0 appear in the retained
run_samplesarrays for the 23active workflows — the pipeline appears to have already filtered zero-AIC runs before
populating the JSON (consistent with the
treated as missing datafiltering logic seenin the debug log for
CI Failure Doctor).History window and date consistency: All 23 active workflows report
history_days: 30and sample dates falling within 2026-07-04–2026-08-03, consistent with the statedwindow; no stale or out-of-range samples were found.
Assumptions
monthperiod, 30-day trailing history, as generated bygh aw forecastat 2026-08-03T10:54:31Z with exit code 0 (no rerun was necessary).
run_samples[].aicvalues recorded in the last 30 days;"projected AIC" = Monte-Carlo-derived P10/P50/P90 extrapolations for the next
week/month, not historical totals.
window (build/lint/CI/test/Dependabot-type jobs) and are excluded from the observed and
projected totals.
their contribution to the aggregate P10/P50/P90 is small (<3% of the monthly P50 total).
Next Actions
P90 ≈ 61,634 as a conservative ceiling for approvals.
needed (check for reduced PR volume or a scheduling gap in that window).
Go Logger Enhancement,Agentic Workflow Audit Agent, andSmoke Copilot— they account for ~59% of observed spend and dominate forecastvariance.
gh aw forecast --evalbacktesting once morehistory accumulates (target ≥10 samples) before relying on their individual P10/P90
ranges for budgeting decisions.