Skip to content

feat(fleet-monitor): summarize token usage by workflow - #343

Merged
don-petry merged 15 commits into
mainfrom
feat/fleet-monitor-token-summary
May 21, 2026
Merged

don-petry merged 15 commits into
mainfrom
feat/fleet-monitor-token-summary

Conversation

@don-petry

@don-petry don-petry commented May 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • add a fleet-monitor step to collect recent token-usage artifacts
  • aggregate JSONL records by workflow (calls, tokens, ET)
  • publish a markdown table to the workflow run summary

Validation

Summary by CodeRabbit

  • Chores
    • Added automated token usage reporting to CI/CD workflow summaries, displaying aggregated metrics including execution time and token consumption across workflows.

Review Change Stack

Copilot AI review requested due to automatic review settings May 21, 2026 12:24
@gemini-code-assist

Copy link
Copy Markdown

Note

Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported.

@coderabbitai

coderabbitai Bot commented May 21, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@don-petry has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 9 minutes and 52 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 20254905-059f-43ac-9915-49977ba7b4f1

📥 Commits

Reviewing files that changed from the base of the PR and between 8683edd and 6ec277a.

📒 Files selected for processing (5)
  • .github/workflows/actions-fleet-monitor.yml
  • .github/workflows/dev-lead-reusable.yml
  • docs/dev-lead/spec.md
  • scripts/engine.sh
  • tests/dev-lead/unit/test_engine_writer.bats
📝 Walkthrough

Walkthrough

This PR adds a new Summarize token usage by workflow step to the actions-fleet-monitor workflow. The step collects token-usage artifacts from GitHub, unpacks their JSONL contents, filters by a configurable lookback period, aggregates token metrics and execution time per workflow, and appends a summary table to the job step summary.

Changes

Token usage monitoring and reporting

Layer / File(s) Summary
Token usage aggregation and reporting step
.github/workflows/actions-fleet-monitor.yml
New health-check step downloads and processes token-usage-* artifacts via GitHub API, filters JSONL records by timestamp using LOOKBACK_DAYS, aggregates token counts (input, cache read, output), call totals, and execution time per workflow with jq, and writes a Markdown summary table or "no artifacts found" message to GITHUB_STEP_SUMMARY.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding a new workflow step to summarize token usage by workflow in the fleet-monitor workflow.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fleet-monitor-token-summary

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Adds a workflow step that collects recent token-usage-* artifacts, aggregates JSONL metrics per workflow, and publishes a markdown summary table to the workflow run summary.

Changes:

  • Download and extract recent token-usage artifacts within a configurable lookback window
  • Aggregate JSONL records by workflow (calls, tokens, elapsed time) into a single summary
  • Emit a markdown table into $GITHUB_STEP_SUMMARY

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread .github/workflows/actions-fleet-monitor.yml Outdated
Comment thread .github/workflows/actions-fleet-monitor.yml
don-petry and others added 2 commits May 21, 2026 07:39
Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@don-petry
don-petry force-pushed the feat/fleet-monitor-token-summary branch from d68423e to 8683edd Compare May 21, 2026 12:39

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/actions-fleet-monitor.yml:
- Around line 122-125: The process substitution feeding the while loop uses `gh
api ... --paginate --jq ...` so failures inside that substitution are swallowed;
change the flow to write the `gh api` output to a temporary file (e.g.,
artifacts_list) using `gh api "repos/$REPO/actions/artifacts" --paginate --jq
'...' > "$artifacts_list"`, check the `gh` command exit code and handle non-zero
by logging an explicit error and exiting, then `while IFS= read -r encoded; do
... done < "$artifacts_list"` to iterate; this ensures `gh api` failures (rate
limits/5xx/bad token) are detected and fail the step instead of producing an
empty loop.
- Around line 105-149: The jq invocation is using the wrong glob
("$jsonl_dir"/*.jsonl) and thus misses files staged under per-artifact subdirs
($jsonl_dir/$id/*.jsonl); replace the single-glob usage at the aggregation step
(the command that writes to the aggregate variable token-usage-summary.json)
with a robust feed from find (e.g., pipe or xargs with -print0/-0) so jq reads
all nested *.jsonl under "$jsonl_dir" (or alternatively enable globstar and use
a recursive glob like "$jsonl_dir"/**/*.jsonl), ensuring the find check and the
jq aggregation operate on the same file set.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 96ad6d02-24a2-4fcb-9419-301c87487f18

📥 Commits

Reviewing files that changed from the base of the PR and between 4303ded and 8683edd.

📒 Files selected for processing (1)
  • .github/workflows/actions-fleet-monitor.yml

Comment thread .github/workflows/actions-fleet-monitor.yml
Comment thread .github/workflows/actions-fleet-monitor.yml Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8683edd31d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread .github/workflows/actions-fleet-monitor.yml
Comment thread .github/workflows/actions-fleet-monitor.yml Outdated
don-petry and others added 3 commits May 21, 2026 07:52
Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
coderabbitai[bot]
coderabbitai Bot previously approved these changes May 21, 2026
donpetry-bot
donpetry-bot previously approved these changes May 21, 2026

@donpetry-bot donpetry-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review — APPROVED ✓

Risk: LOW
Reviewed commit: 646ba3fe803328de05221fad6336b3ee7112e050
Review mode: triage-approved (single reviewer)

Summary

Confirming the triage assessment. This PR has two unrelated but small changes: (1) a new Summarize token usage by workflow step in actions-fleet-monitor.yml that downloads recent token-usage-* artifacts, aggregates JSONL records per workflow with jq, and writes a Markdown table to $GITHUB_STEP_SUMMARY; (2) a gemini model-name bump (1.5-pro → 2.5-pro) in scripts/engine.sh and matching doc table in docs/dev-lead/spec.md to fix ModelNotFound during rate-limit fallback. Net diff is +93/−9 across 3 files.

Linked issue analysis

No linked issue. The PR description is self-explanatory and the change matches it. CodeRabbit's pre-merge checks (title, description, scope, docstring coverage) all passed.

Findings

  • Workflow step hygiene — uses set -euo pipefail, mktemp -d with trap cleanup, scoped GH_TOKEN from secrets.GH_PAT_WORKFLOWS (same secret already used by the prior fleet_monitor.sh step). No new permissions or secret surface.
  • Prior reviewer concerns resolved — CodeRabbit's earlier CHANGES_REQUESTED flagged two issues that are both fixed in this head SHA:
    • gh api output is now captured to $artifacts_file and iterated via while … < "$artifacts_file" (no process substitution swallowing failures).
    • JSONL files are flat-copied into $jsonl_dir with ${id}- prefix, so the jq -s … "$jsonl_dir"/*.jsonl glob matches what the preceding find check verifies exists. CodeRabbit's most recent review (same SHA) is APPROVED.
  • Engine model bump — gemini-1.5-pro → gemini-2.5-pro in ENGINE_DEEP_MODEL, ENGINE_AUDIT_MODEL, ENGINE_ACTION_MODEL, ENGINE_SINGLE_MODEL, with matching ENGINE_LABEL strings and the docs/dev-lead/spec.md table. Doc and code are in sync. Triage model (gemini-2.0-flash) intentionally unchanged.
  • Minor / non-blocking — unzip -o extracts artifacts into per-id directories without path-traversal hardening; acceptable for a scheduled monitor consuming artifacts the org itself produces, but worth noting for defence-in-depth if this code is ever generalised. No action required.

CI status

All required checks green: CI (Lint, ShellCheck, Agent Security Scan, Compile agentic workflows, Secret scan/gitleaks), CodeQL (actions), SonarCloud (Quality Gate passed, 0 new issues), Tests (unit-tests, bats), Lint (shellcheck, gh-aw-compile, validate-agent-profiles), Dependency audit, AgentShield, PR Review Agent. mergeStateStatus: CLEAN, mergeable: MERGEABLE.


Reviewed automatically by the PR-review agent (single-reviewer mode: opus 4.7). Reply if you need a human review.

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@don-petry
don-petry dismissed stale reviews from donpetry-bot and coderabbitai[bot] via 6aaee06 May 21, 2026 20:29
don-petry and others added 2 commits May 21, 2026 15:32
Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — rate-limited (intent: human-pr)

PR: #343
The retry cron will re-attempt automatically.

@don-petry

Copy link
Copy Markdown
Collaborator Author

Note

I received your request but all AI engines are currently rate-limited. I'll retry automatically once the rate limit clears.
Rate limit resets at: unknown

@sonarqubecloud

Copy link
Copy Markdown

don-petry added a commit that referenced this pull request Jun 23, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 23, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 2, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 2, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 8, 2026
* feat(fleet-monitor): add token usage summary by workflow

Download recent token-usage artifacts, aggregate JSONL records by workflow, and publish token totals and ET in the workflow step summary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token summary aggregation

Avoid JSONL filename collisions by staging per-artifact directories and sort by workflow before jq group_by to ensure correct totals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token artifact aggregation file matching

Prefix copied JSONL filenames with artifact IDs to avoid collisions while keeping flat aggregation input for jq.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fleet monitor token artifact collection robustness

Use explicit artifact-list capture (no process substitution masking), and pin artifact source repo to .github-private for workflow_call correctness.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix gemini fallback model for dev-lead

Switch gemini deep/audit/action/single models to gemini-2.5-pro to avoid ModelNotFound failures during rate-limit fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix dev-lead auto-merge guard for skip intent

Require INTENT_PR_NUMBER for enable-auto-merge step so self-review skip events don't invoke fix-reviews with empty PR context.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback detection for stderr rate-limit errors

Capture stderr in run_writer token logs so quota/rate-limit responses from Gemini and Copilot are classified as retryable (exit 2) and can fall through to next engine; add regression test.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix fallback when copilot token is classic PAT

Skip copilot fallback when COPILOT_GITHUB_TOKEN is a classic ghp token (unsupported by Copilot CLI), preserving rate-limit semantics and preventing hard failures after claude/gemini exhaustion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants