Skip to content

feat(token-report): org-wide weekly Token Cost Observatory report - #456

Merged
don-petry merged 12 commits into
mainfrom
feat/token-report-weekly
Jun 7, 2026
Merged

don-petry merged 12 commits into
mainfrom
feat/token-report-weekly

Conversation

@don-petry

@don-petry don-petry commented Jun 7, 2026 •

Copy link
Copy Markdown
Collaborator

Why

We logged discussion #332 → issue #333 about token metrics, which shipped per-call JSONL logging (PR #334) and a summary step (PR #343). But the report was never actually visible:

  1. Buried — the only output was the daily fleet-monitor Step Summary; no issue, no notification, no persisted doc. Phase 4 (weekly report to an issue) was a stretch goal that never landed.
  2. ~6% coverage — the summary hardcoded REPO: petry-projects/.github-private. The agents run as reusable workflows in each caller repo, so their token-usage-* artifacts land there. The summary saw 160 artifacts and was blind to ~2,300 others (markets, google-app-scripts, bmad-bgreat-suite, …).

What

  • scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage-* artifact in the lookback window, and renders an Effective Tokens (ET) rollup by workflow/tier/model and by repository. Pure render_*/aggregate_* functions (unit-tested); main() does the network I/O.
  • .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) + workflow_dispatch. Posts the report as a comment on a single pinned tracking issue (label token-report).
  • actions-fleet-monitor.yml — the inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug; net −82 lines).
  • Tests (tests/token_report.bats + fixtures), docs (docs/token-report.md).

Sample output (live, 2-day window)

## Top cost drivers (workflow / tier / model)
| pr-review | single | claude-opus-4-7   | 16 calls | ET 1,209,180 | 52% |
| dev-lead  | writer | claude-sonnet-4-6 |106 calls | ET   959,304 | 41% |
| pr-review | triage | claude-haiku-4-5  | 21 calls | ET    97,683 |  4% |

The opus-4.7 pr-review tier dominates (15× multiplier × 4× output weight) — exactly the optimization signal the observatory was meant to surface.

Test plan

  • bats tests/token_report.bats (11 tests) + tests/fleet_report.bats
  • shellcheck --severity=warning on all three scripts
  • End-to-end run against the live org (8 repos, 127 runs, 148 calls)
  • markdownlint + YAML validity on changed files

Closes #206.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Weekly organization-wide token cost reports now automatically generate and post to a tracking issue
    • Manual token report generation available via workflow dispatch
    • Token usage aggregation by workflow and repository with top cost drivers identified
  • Documentation

    • Added comprehensive token cost reporting documentation
  • Tests

    • Added test suite for token cost reporting functionality

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings June 7, 2026 03:21
@don-petry
don-petry requested a review from a team as a code owner June 7, 2026 03:21
@coderabbitai

coderabbitai Bot commented Jun 7, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@don-petry, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 31 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 23468977-3b01-4df0-a68d-151101359152

📥 Commits

Reviewing files that changed from the base of the PR and between b96f7b8 and 67b6b95.

⛔ Files ignored due to path filters (1)
  • scripts/lib/model-pricing.tsv is excluded by !**/*.tsv
📒 Files selected for processing (13)
  • .github/workflows/lint.yml
  • AGENTS.md
  • docs/token-report.md
  • scripts/engine.sh
  • scripts/lib/model-pricing.sh
  • scripts/lib/token-metrics.sh
  • scripts/token_report.sh
  • tests/dev-lead/fixtures/engines/stub-claude
  • tests/dev-lead/fixtures/engines/stub-gemini
  • tests/dev-lead/unit/test_engine_writer.bats
  • tests/dev-lead/unit/test_token_metrics.bats
  • tests/model_pricing.bats
  • tests/token_report.bats
📝 Walkthrough

Walkthrough

This PR introduces an org-wide token cost reporting system that collects token-usage-* artifacts from all non-archived repositories, aggregates them by workflow/tier/model and by repository, and delivers weekly Markdown reports to a tracking issue. The fleet monitor workflow is updated to display org-wide totals, and comprehensive testing validates all script functions.

Changes

Token Cost Observatory

Layer / File(s) Summary
Token report collection, aggregation, and rendering script
scripts/token_report.sh
collect_org_jsonl() discovers non-archived org repos and downloads matching token-usage-* artifact ZIPs via GitHub API, extracts JSONL token records, tags them with repo origin, and returns counts. Helper functions _fmt_int() format numbers with thousands separators, aggregate_by_workflow() and aggregate_by_repo() group records and sum call/ET totals by jq, and render_token_report() produces Markdown sections for totals, top cost drivers (by workflow/tier/model), and per-repository breakdown. _extract_zip() extracts ZIPs via unzip or python3 fallback. main() orchestrates collection and rendering, outputs to GITHUB_STEP_SUMMARY, and gates execution via entrypoint guard.
Test suite and fixtures for token report script
tests/token_report.bats, tests/fixtures/token_jsonl/run-a.jsonl, tests/fixtures/token_jsonl/run-b.jsonl
Bats test suite validates _fmt_int() formatting, aggregate_by_workflow() grouping/summing and ET sorting, aggregate_by_repo() grouping and ET sorting, and render_token_report() Markdown sections, totals, and ET percentage shares. Tests include empty-directory handling and error/resilience cases for collect_org_jsonl() (stubbed gh API failures and partial failures). JSONL fixtures contain token metadata (timestamps, workflow/tier, model, token counts, elapsed time, PR context) for runs a and b across multiple workflows.
Token Cost Report workflow with tracking-issue posting
.github/workflows/token-report.yml
Schedules weekly (Mondays 08:00 UTC) and accepts manual dispatch with org and lookback-window inputs. Runs scripts/token_report.sh to collect metrics, appends output to Step Summary. Conditional github-script step posts results to a tracking issue: reads token_report.md, fails if empty, truncates to 65K characters with run-summary link if needed, ensures token-report label exists, finds or creates open tracking issue by pinned title match, and posts report as comment.
Fleet Monitor workflow integration of org-wide token reporting
.github/workflows/actions-fleet-monitor.yml
Removes inline "Summarize token usage by workflow" step that enumerated and parsed token-usage-* artifacts. Replaces with "Summarize token usage (org-wide)" step that runs scripts/token_report.sh to aggregate org-wide totals, appends output to Step Summary, and is non-blocking via continue-on-error: true.
CI linting and testing updates
.github/workflows/lint.yml
Shellcheck step adds scripts/token_report.sh to linting targets. Bats test step adds tests/token_report.bats to test execution.
Token Cost Observatory documentation
docs/token-report.md, docs/actions-fleet-monitor.md
docs/token-report.md describes the org-wide "Token Cost Observatory — Weekly Report" including ET cost normalization with model multipliers and weighting, delivery via weekly pinned tracking-issue comments and fleet-monitor step summaries, usage modes (scheduled weekly, manual gh workflow run dispatch, local script execution), architecture mapping to scripts/tests/workflows, required environment variables, and stated limitations (artifact retention, cache-token visibility, per-repo artifact listing scalability). docs/actions-fleet-monitor.md adds documentation of the org-wide token rollup appended to Step Summary and links to the token-report documentation.

Sequence Diagram(s)

sequenceDiagram
  participant main as main()
  participant collect as collect_org_jsonl
  participant gh_api as GitHub API
  participant extract as _extract_zip
  participant agg_wf as aggregate_by_workflow
  participant agg_repo as aggregate_by_repo
  participant render as render_token_report
  main->>collect: create temp dir, call collect
  collect->>gh_api: list non-archived org repos
  gh_api-->>collect: repo list
  collect->>gh_api: per-repo, list token-usage-* artifacts
  gh_api-->>collect: artifact metadata
  collect->>collect: download artifact ZIPs
  collect->>extract: extract JSONL files
  extract-->>collect: extracted JSONL records
  collect->>collect: tag records with repo, write to dir
  collect-->>main: repo_count, artifact_count
  main->>agg_wf: aggregate_by_workflow
  agg_wf-->>main: JSON workflow totals
  main->>agg_repo: aggregate_by_repo
  agg_repo-->>main: JSON repo totals
  main->>render: render with both aggregations
  render-->>main: Markdown report
  main->>main: output to GITHUB_STEP_SUMMARY
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

The PR introduces a new token reporting system with moderate scope: the core script (scripts/token_report.sh) contains ~280 lines of logic with multiple functions (collection, aggregation, rendering), a comprehensive test suite with ~130 lines covering unit and integration scenarios, a new workflow file with ~120 lines including complex github-script logic, and scattered updates to existing workflows and documentation. The changes are coherent (all serve the token reporting feature) but heterogeneous enough to require separate reasoning for the script logic, test coverage, workflow integration, and CI updates. The script contains medium-density logic (jq aggregations, ZIP extraction with fallbacks, GitHub API pagination), and the github-script posting logic is moderately complex (label/issue management, truncation).

Possibly related PRs

  • petry-projects/.github-private#334: The main PR's scripts/token_report.sh aggregates and downloads token-usage-* JSONL artifacts produced by the retrieved PR's scripts/lib/token-metrics.sh instrumentation in scripts/engine.sh.
  • petry-projects/.github-private#343: The main PR removes and replaces the "Summarize token usage by workflow" step added in PR #343 with scripts/token_report.sh org-wide aggregation.
  • petry-projects/.github-private#197: Main PR extends Fleet Monitor work by adding token reporting and updating shared .github/workflows/lint.yml and docs/actions-fleet-monitor.md touchpoints.
🚥 Pre-merge checks | ✅ 3 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The PR implements token tracking and reporting infrastructure but does not address the proactive provider headroom check acceptance criteria from issue #206 (rate-limit detection, skip logic, DEV_LEAD_USAGE_THRESHOLD env var). This PR appears to be foundational work for future rate-limit checking. Address issue #206's acceptance criteria in a follow-up PR: implement rate-limited marker detection in engine invocation logic with skip/re-queue behavior.
Docstring Coverage ⚠️ Warning Docstring coverage is 75.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the main change: introduction of an org-wide weekly token cost reporting system, accurately reflecting the PR's primary objectives.
Out of Scope Changes check ✅ Passed All changes are directly related to building the token cost reporting system (scripts, workflows, tests, docs) needed to track provider usage, which is foundational to the headroom-check feature.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/token-report-weekly

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — review-changes (no-changes)

No changes were needed for this PR.

@don-petry
don-petry enabled auto-merge (squash) June 7, 2026 03:22

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces an org-wide Token Cost Observatory report, adding scripts/token_report.sh to aggregate LLM token spend across all non-archived repositories, along with corresponding documentation and BATS unit tests. The review feedback highlights critical reliability and portability issues in the script: first, glob expansion failures when no .jsonl files are present will cause the script to crash under set -euo pipefail in both aggregate_by_workflow and aggregate_by_repo; second, the use of GNU-specific date -u -d syntax will cause failures on macOS (BSD), which can be resolved by implementing a portable date fallback and filtering artifacts directly within jq.

Comment thread scripts/token_report.sh Outdated
Comment thread scripts/token_report.sh Outdated
Comment thread scripts/token_report.sh Outdated
donpetry-bot
donpetry-bot previously approved these changes Jun 7, 2026

@donpetry-bot donpetry-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review — APPROVED ✓

Risk: LOW
Reviewed commit: b5215c648bfea357cdde4ec8fada909100ea285d
Review mode: triage-approved (single reviewer)

Summary

This PR delivers Phase 4 of the Token Cost Observatory (#333): an org-wide weekly token-cost report. The implementation cleanly separates pure rendering/aggregation (unit-tested) from network I/O, replaces the hardcoded single-repo inline summary in actions-fleet-monitor.yml with a call to the shared script (net −82 lines), and adds a weekly cron workflow that posts to a single pinned tracking issue. Triage's low-risk assessment holds up on confirmation review.

Linked issue analysis

The PR body references the Token Cost Observatory chain (discussion #332 → issue #333 → PRs #334/#343) and describes this work as the unlanded Phase 4 weekly-report stretch goal from #333. That framing is accurate — scripts/token_report.sh and .github/workflows/token-report.yml directly implement the “Phase 4 — Weekly Report (stretch)” section of #333.

Note (non-blocking): the Closes #206 trailer points to a different feature — issue #206 is feat(dev-lead): proactive provider headroom check before engine invocation, which is about pre-flight rate-limit detection, not a token-cost report. The acceptance criteria there (session-scoped rate-limit markers, fail-open invocation gates, DEV_LEAD_USAGE_THRESHOLD) are not addressed by this PR. Worth confirming before merge whether the trailer should be Closes #333 (already closed) or a new tracking issue for Phase 4, and whether #206 should remain open for the proactive-headroom work.

Findings

Code quality — strong

  • scripts/token_report.sh cleanly separates concerns: pure aggregate_by_workflow, aggregate_by_repo, render_token_report, and _fmt_int are unit-tested; main/collect_org_jsonl handle network I/O. The BASH_SOURCE[0] = $0 guard correctly allows sourcing from bats without executing main.
  • set -euo pipefail, mktemp + trap … EXIT/RETURN for cleanup, and 2>/dev/null || true on fallible discovery calls so a single bad repo doesn't kill the whole run.
  • The unzip-with-Python-fallback path is a nice touch for minimal environments, though CI runners always have unzip.
  • 11 bats tests covering formatting, grouping, ET sort order, totals, percentage share, and the empty-dir path. Fixture files are realistic JSONL records with a repo field.

Workflow / security

  • actions/checkout and actions/github-script are pinned to commit SHAs (v6.0.3, v9.0.0) — matches repo convention.
  • permissions: contents: read, issues: write — minimal and correct for the comment-posting step.
  • GH_PAT_WORKFLOWS (cross-org actions:read) is required to read other repos' artifacts and is correctly scoped to the collection step; the comment-posting step uses github.token (in-repo only). The header comment in token-report.yml explains the split clearly.
  • concurrency: token-report-${{ inputs.org || 'petry-projects' }} with cancel-in-progress: true is appropriate for a weekly cron.
  • The 65 000-char MAX truncation guard with a link-back to the run summary handles oversized reports defensively.
  • No shell injection surface: all interpolated values ($repo, $id, $created_at) come from gh api JSON and are wrapped in printf '%s' / safe quoting. date -u -d "$created_at" has a || echo 0 fallback.
  • Find-or-create issue logic correctly paginates listForRepo and matches on exact title.

Minor observations (non-blocking)

  • collect_org_jsonl uses gh api "orgs/${ORG}/repos?per_page=100&type=all" --paginate. For very large orgs this is O(repos) sequential artifact-list calls, as the docs already note — fine for current scale.
  • The _extract_zip Python fallback is dead code on GitHub-hosted runners; harmless and cheap to keep.

CI status

All 27 check runs are SUCCESS or SKIPPED:

  • AgentShield ✓
  • CodeQL (actions) ✓
  • SonarCloud — Quality Gate passed (0 new issues, 0 hotspots) ✓
  • Lint: shellcheck (now covers token_report.sh), bats (now runs tests/token_report.bats), validate-agent-profiles, gh-aw-compile ✓
  • CI: Lint, ShellCheck, Compile agentic workflows, Secret scan (gitleaks), Agent Security Scan ✓
  • Tests: unit-tests ✓
  • Dependency audit ✓
  • Dev-Lead Agent dispatch ✓
  • CodeRabbit and SonarCloud Code Analysis ✓

CodeRabbit hit a per-org rate limit and didn't post a substantive review, but its status is SUCCESS and the other static-analysis layers (CodeQL, SonarCloud, ShellCheck, gitleaks, AgentShield) all cleared.

mergeStateStatus: BLOCKED is solely due to required review, not failing checks.


Reviewed automatically by the PR-review agent (single-reviewer mode: opus 4.7). Reply if you need a human review.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds an org-wide “Token Cost Observatory” weekly report to make token-usage artifacts visible and actionable across all non-archived repos in the org (instead of being limited to .github-private), and reuses the same reporting logic in the daily fleet monitor Step Summary.

Changes:

  • Introduces scripts/token_report.sh to collect token-usage-* artifacts org-wide and render aggregated Markdown (by workflow/tier/model and by repo).
  • Adds a scheduled workflow (.github/workflows/token-report.yml) to post the weekly report as a comment on a tracking issue labeled token-report.
  • Updates the fleet monitor workflow to reuse the shared report script and extends CI to shellcheck + test it.

Reviewed changes

Copilot reviewed 9 out of 9 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
scripts/token_report.sh New org-wide artifact collector + Markdown report renderer used by both weekly report and fleet-monitor summary.
tests/token_report.bats Unit tests for the pure aggregation/rendering functions (no network I/O).
tests/fixtures/token_jsonl/run-a.jsonl JSONL fixture data for aggregation/rendering tests.
tests/fixtures/token_jsonl/run-b.jsonl Additional JSONL fixture data for aggregation/rendering tests.
.github/workflows/token-report.yml Weekly cron + manual dispatch workflow to generate and post the report to a tracking issue.
.github/workflows/actions-fleet-monitor.yml Replaces single-repo token summary step with org-wide report via scripts/token_report.sh.
.github/workflows/lint.yml Adds shellcheck coverage for the new script and runs the new bats tests.
docs/token-report.md Documentation for the weekly report, ET metric, and operational usage.
docs/actions-fleet-monitor.md Documents that fleet-monitor now includes an org-wide token usage rollup.

Comment thread scripts/token_report.sh Outdated
Comment thread scripts/token_report.sh Outdated
Comment thread scripts/token_report.sh Outdated
Comment thread tests/fixtures/token_jsonl/run-b.jsonl Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b5215c648b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/token_report.sh Outdated
Comment thread .github/workflows/actions-fleet-monitor.yml Outdated
Comment thread scripts/token_report.sh Outdated
coderabbitai[bot]
coderabbitai Bot previously approved these changes Jun 7, 2026
@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — fix-reviews (applied)

Changes committed and pushed.

@don-petry
don-petry dismissed stale reviews from coderabbitai[bot] and donpetry-bot via e69063f June 7, 2026 03:28

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e69063f731

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/token_report.sh Outdated
coderabbitai[bot]
coderabbitai Bot previously approved these changes Jun 7, 2026
@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — review-changes (applied)

Changes committed and pushed.

@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — review-changes (applied)

Changes committed and pushed.

@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — rate-limited (intent: fix-reviews)

PR: #456
The retry cron will re-attempt automatically. Rate limit resets at: 2026-06-07T04:23:06Z

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b96f7b8fa8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/token_report.sh Outdated
Comment thread .github/workflows/token-report.yml

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/lint.yml:
- Line 25: Update the shellcheck invocation under the run step that currently
reads "shellcheck --severity=warning scripts/fleet_monitor.sh
scripts/fleet_report.sh scripts/token_report.sh" to include explicit Bash mode
by adding the flag "--shell=bash" so the command becomes "shellcheck
--severity=warning --shell=bash ..." ensuring linting uses Bash semantics for
scripts/fleet_monitor.sh, scripts/fleet_report.sh, and scripts/token_report.sh.

In @.github/workflows/token-report.yml:
- Around line 51-57: The checkout step using
actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 with token: ${{
secrets.GH_PAT_WORKFLOWS }} persists the PAT in the repo's git config; update
the checkout "with" block for that step (the Checkout agent repo step) to set
persist-credentials: false so the PAT is not stored in the local git credential
helper after checkout.
- Around line 34-37: Remove the workflow-wide "issues: write" permission from
the top-level permissions block and instead grant "issues: write" only to the
specific job named "report" by adding a permissions block under jobs.report
(keep top-level permissions as minimal as required, e.g., contents: read) so
only the report job has write access to issues.

In `@scripts/token_report.sh`:
- Around line 233-235: The artifact download/extraction loop currently swallows
errors with bare `continue`, causing silent undercounts; update the failure
handlers for the `gh api "repos/${repo}/actions/artifacts/${id}/zip"` download
and the `_extract_zip "$zip" "$ex"` extraction to emit descriptive warnings to
stderr that include the repo, artifact id, and target paths (e.g., zip and ex)
and describe which step failed, then continue; reference the `gh api`
invocation, the `_extract_zip` call, and variables `repo`, `id`, `zip`, and `ex`
to locate where to add the warning messages.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 1266723d-8f5e-4314-876a-1da1272cf9d3

📥 Commits

Reviewing files that changed from the base of the PR and between 5a9a247 and b96f7b8.

📒 Files selected for processing (9)
  • .github/workflows/actions-fleet-monitor.yml
  • .github/workflows/lint.yml
  • .github/workflows/token-report.yml
  • docs/actions-fleet-monitor.md
  • docs/token-report.md
  • scripts/token_report.sh
  • tests/fixtures/token_jsonl/run-a.jsonl
  • tests/fixtures/token_jsonl/run-b.jsonl
  • tests/token_report.bats

Comment thread .github/workflows/lint.yml Outdated
Comment thread .github/workflows/token-report.yml
Comment thread .github/workflows/token-report.yml
Comment thread scripts/token_report.sh Outdated
@don-petry

Copy link
Copy Markdown
Collaborator Author

@coderabbitai resolve

@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — fix-reviews (applied)

Changes committed and pushed.

don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jul 14, 2026
…aily-digest accumulation (#1203) (#1204)

* fix(fleet-monitor): auto-close resolved fleet-tracker issues (#1203)

The monitor creates/updates a fleet-tracker issue per (repo × workflow) above
10% failure, but never closed them when the workflow recovered or was deleted —
so 107 stale alerts accumulated (some with 'Last updated' stamps ~55 days old
while the monitor runs daily; 4 for the deleted claude.yml #456).

Add a step that closes any open fleet-tracker issue whose 'Last updated on <date>'
stamp is older than STALE_DAYS (default 3 = 3 missed daily runs), with a
✅ Auto-resolved comment. It reopens automatically (fresh issue) if the workflow
crosses the threshold again — same lifecycle as org-scorecard.

Safety: gated on fleet_high_failure.json existing (a failed scan can't mass-close);
staleness uses each issue's own stamp, not the current run, so a single bad-scan
day never false-closes; comment + close are each guarded so a comment-capped issue
still closes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Juznz5V6su81ffSND8fg7s

* fix(reviews): address review comments [skip ci-relay]

* fix(fleet-monitor): close prior daily digests instead of accumulating them

The 'workflow failures detected <date>' digest embeds the date in its title, so
issues.create ran unconditionally every day — 42 open digests had piled up. Close
the prior open digests (health-check label, matching title prefix) before opening
today's, so only the current snapshot stays open. Complements the fleet-tracker
auto-close in this PR; both stop the monitor accumulating stale issues.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Juznz5V6su81ffSND8fg7s

* chore: dev-lead update (review-changes) [skip ci-relay]

* fix(bot): address bot feedback [skip ci-relay]

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 2, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 2, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
…y) (#460)

* fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency)

review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time.
$$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$`
sidecar collided between the two concurrent tier-2 calls — one engine's usage could
overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID
doesn't work either: it differs between the chain's pipe subshell and the reader.)

Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path
(unique by construction) and inherited by the engine's pipeline subshell, so writer
and reader agree while concurrent calls never share a file. $$ remains a fallback for
non-concurrent direct callers.

- token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT.
- engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key.
- tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key
  and fallback assertions; correct the misleading "$$ isolates parallel" test.
- token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on
  NUL bytes; use truncated JSON, which jq rejects across versions).

Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* test(token-metrics): cover mktemp-failure fallback for per-call usage key

Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp
(PR #460 review, copilot-pull-request-reviewer ×3):

- unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the
  mktemp-failure state), so it never reuses a prior/inherited key.
- engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT
  reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(dev-lead): proactive provider headroom check before engine invocation

3 participants