Repository navigation
feat(fleet-monitor): summarize token usage by workflow - #343
Conversation
|
Note Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported. |
|
Warning Rate limit exceeded
You’ve run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (5)
📝 WalkthroughWalkthroughThis PR adds a new ChangesToken usage monitoring and reporting
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Pull request overview
Note
Copilot was unable to run its full agentic suite in this review.
Adds a workflow step that collects recent token-usage-* artifacts, aggregates JSONL metrics per workflow, and publishes a markdown summary table to the workflow run summary.
Changes:
- Download and extract recent token-usage artifacts within a configurable lookback window
- Aggregate JSONL records by workflow (calls, tokens, elapsed time) into a single summary
- Emit a markdown table into
$GITHUB_STEP_SUMMARY
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
d68423e to
8683edd
Compare
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/actions-fleet-monitor.yml:
- Around line 122-125: The process substitution feeding the while loop uses `gh
api ... --paginate --jq ...` so failures inside that substitution are swallowed;
change the flow to write the `gh api` output to a temporary file (e.g.,
artifacts_list) using `gh api "repos/$REPO/actions/artifacts" --paginate --jq
'...' > "$artifacts_list"`, check the `gh` command exit code and handle non-zero
by logging an explicit error and exiting, then `while IFS= read -r encoded; do
... done < "$artifacts_list"` to iterate; this ensures `gh api` failures (rate
limits/5xx/bad token) are detected and fail the step instead of producing an
empty loop.
- Around line 105-149: The jq invocation is using the wrong glob
("$jsonl_dir"/*.jsonl) and thus misses files staged under per-artifact subdirs
($jsonl_dir/$id/*.jsonl); replace the single-glob usage at the aggregation step
(the command that writes to the aggregate variable token-usage-summary.json)
with a robust feed from find (e.g., pipe or xargs with -print0/-0) so jq reads
all nested *.jsonl under "$jsonl_dir" (or alternatively enable globstar and use
a recursive glob like "$jsonl_dir"/**/*.jsonl), ensuring the find check and the
jq aggregation operate on the same file set.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 96ad6d02-24a2-4fcb-9419-301c87487f18
📒 Files selected for processing (1)
.github/workflows/actions-fleet-monitor.yml
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8683edd31d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
donpetry-bot
left a comment
There was a problem hiding this comment.
Automated review — APPROVED ✓
Risk: LOW
Reviewed commit: 646ba3fe803328de05221fad6336b3ee7112e050
Review mode: triage-approved (single reviewer)
Summary
Confirming the triage assessment. This PR has two unrelated but small changes: (1) a new Summarize token usage by workflow step in actions-fleet-monitor.yml that downloads recent token-usage-* artifacts, aggregates JSONL records per workflow with jq, and writes a Markdown table to $GITHUB_STEP_SUMMARY; (2) a gemini model-name bump (1.5-pro → 2.5-pro) in scripts/engine.sh and matching doc table in docs/dev-lead/spec.md to fix ModelNotFound during rate-limit fallback. Net diff is +93/−9 across 3 files.
Linked issue analysis
No linked issue. The PR description is self-explanatory and the change matches it. CodeRabbit's pre-merge checks (title, description, scope, docstring coverage) all passed.
Findings
- Workflow step hygiene — uses
set -euo pipefail,mktemp -dwithtrapcleanup, scopedGH_TOKENfromsecrets.GH_PAT_WORKFLOWS(same secret already used by the priorfleet_monitor.shstep). No new permissions or secret surface. - Prior reviewer concerns resolved — CodeRabbit's earlier
CHANGES_REQUESTEDflagged two issues that are both fixed in this head SHA:gh apioutput is now captured to$artifacts_fileand iterated viawhile … < "$artifacts_file"(no process substitution swallowing failures).- JSONL files are flat-copied into
$jsonl_dirwith${id}-prefix, so thejq -s … "$jsonl_dir"/*.jsonlglob matches what the precedingfindcheck verifies exists. CodeRabbit's most recent review (same SHA) is APPROVED.
- Engine model bump —
gemini-1.5-pro→gemini-2.5-proinENGINE_DEEP_MODEL,ENGINE_AUDIT_MODEL,ENGINE_ACTION_MODEL,ENGINE_SINGLE_MODEL, with matchingENGINE_LABELstrings and thedocs/dev-lead/spec.mdtable. Doc and code are in sync. Triage model (gemini-2.0-flash) intentionally unchanged. - Minor / non-blocking —
unzip -oextracts artifacts into per-id directories without path-traversal hardening; acceptable for a scheduled monitor consuming artifacts the org itself produces, but worth noting for defence-in-depth if this code is ever generalised. No action required.
CI status
All required checks green: CI (Lint, ShellCheck, Agent Security Scan, Compile agentic workflows, Secret scan/gitleaks), CodeQL (actions), SonarCloud (Quality Gate passed, 0 new issues), Tests (unit-tests, bats), Lint (shellcheck, gh-aw-compile, validate-agent-profiles), Dependency audit, AgentShield, PR Review Agent. mergeStateStatus: CLEAN, mergeable: MERGEABLE.
Reviewed automatically by the PR-review agent (single-reviewer mode: opus 4.7). Reply if you need a human review.
Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
6aaee06
Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Dev-Lead — rate-limited (intent: human-pr)PR: #343 |
|
Note I received your request but all AI engines are currently rate-limited. I'll retry automatically once the rate limit clears. |
|
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(fleet-monitor): add token usage summary by workflow Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token summary aggregation Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix token artifact aggregation file matching Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fleet monitor token artifact collection robustness Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix gemini fallback model for dev-lead Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix dev-lead auto-merge guard for skip intent Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback detection for stderr rate-limit errors Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix fallback when copilot token is classic PAT Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>



Summary
Validation
Summary by CodeRabbit