Repository navigation
feat(token-report): org-wide weekly Token Cost Observatory report - #456
Conversation
The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
Warning Review limit reached
More reviews will be available in 3 minutes and 31 seconds. Learn how PR review limits work. Your organization has run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (13)
📝 WalkthroughWalkthroughThis PR introduces an org-wide token cost reporting system that collects ChangesToken Cost Observatory
Sequence Diagram(s)sequenceDiagram
participant main as main()
participant collect as collect_org_jsonl
participant gh_api as GitHub API
participant extract as _extract_zip
participant agg_wf as aggregate_by_workflow
participant agg_repo as aggregate_by_repo
participant render as render_token_report
main->>collect: create temp dir, call collect
collect->>gh_api: list non-archived org repos
gh_api-->>collect: repo list
collect->>gh_api: per-repo, list token-usage-* artifacts
gh_api-->>collect: artifact metadata
collect->>collect: download artifact ZIPs
collect->>extract: extract JSONL files
extract-->>collect: extracted JSONL records
collect->>collect: tag records with repo, write to dir
collect-->>main: repo_count, artifact_count
main->>agg_wf: aggregate_by_workflow
agg_wf-->>main: JSON workflow totals
main->>agg_repo: aggregate_by_repo
agg_repo-->>main: JSON repo totals
main->>render: render with both aggregations
render-->>main: Markdown report
main->>main: output to GITHUB_STEP_SUMMARY
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes The PR introduces a new token reporting system with moderate scope: the core script ( Possibly related PRs
🚥 Pre-merge checks | ✅ 3 | ❌ 2❌ Failed checks (2 warnings)
✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Dev-Lead — review-changes (no-changes)No changes were needed for this PR. |
There was a problem hiding this comment.
Code Review
This pull request introduces an org-wide Token Cost Observatory report, adding scripts/token_report.sh to aggregate LLM token spend across all non-archived repositories, along with corresponding documentation and BATS unit tests. The review feedback highlights critical reliability and portability issues in the script: first, glob expansion failures when no .jsonl files are present will cause the script to crash under set -euo pipefail in both aggregate_by_workflow and aggregate_by_repo; second, the use of GNU-specific date -u -d syntax will cause failures on macOS (BSD), which can be resolved by implementing a portable date fallback and filtering artifacts directly within jq.
donpetry-bot
left a comment
There was a problem hiding this comment.
Automated review — APPROVED ✓
Risk: LOW
Reviewed commit: b5215c648bfea357cdde4ec8fada909100ea285d
Review mode: triage-approved (single reviewer)
Summary
This PR delivers Phase 4 of the Token Cost Observatory (#333): an org-wide weekly token-cost report. The implementation cleanly separates pure rendering/aggregation (unit-tested) from network I/O, replaces the hardcoded single-repo inline summary in actions-fleet-monitor.yml with a call to the shared script (net −82 lines), and adds a weekly cron workflow that posts to a single pinned tracking issue. Triage's low-risk assessment holds up on confirmation review.
Linked issue analysis
The PR body references the Token Cost Observatory chain (discussion #332 → issue #333 → PRs #334/#343) and describes this work as the unlanded Phase 4 weekly-report stretch goal from #333. That framing is accurate — scripts/token_report.sh and .github/workflows/token-report.yml directly implement the “Phase 4 — Weekly Report (stretch)” section of #333.
Note (non-blocking): the Closes #206 trailer points to a different feature — issue #206 is feat(dev-lead): proactive provider headroom check before engine invocation, which is about pre-flight rate-limit detection, not a token-cost report. The acceptance criteria there (session-scoped rate-limit markers, fail-open invocation gates, DEV_LEAD_USAGE_THRESHOLD) are not addressed by this PR. Worth confirming before merge whether the trailer should be Closes #333 (already closed) or a new tracking issue for Phase 4, and whether #206 should remain open for the proactive-headroom work.
Findings
Code quality — strong
scripts/token_report.shcleanly separates concerns: pureaggregate_by_workflow,aggregate_by_repo,render_token_report, and_fmt_intare unit-tested;main/collect_org_jsonlhandle network I/O. TheBASH_SOURCE[0] = $0guard correctly allows sourcing from bats without executingmain.set -euo pipefail, mktemp +trap … EXIT/RETURNfor cleanup, and2>/dev/null || trueon fallible discovery calls so a single bad repo doesn't kill the whole run.- The
unzip-with-Python-fallback path is a nice touch for minimal environments, though CI runners always haveunzip. - 11 bats tests covering formatting, grouping, ET sort order, totals, percentage share, and the empty-dir path. Fixture files are realistic JSONL records with a
repofield.
Workflow / security
actions/checkoutandactions/github-scriptare pinned to commit SHAs (v6.0.3, v9.0.0) — matches repo convention.permissions: contents: read, issues: write— minimal and correct for the comment-posting step.GH_PAT_WORKFLOWS(cross-orgactions:read) is required to read other repos' artifacts and is correctly scoped to the collection step; the comment-posting step usesgithub.token(in-repo only). The header comment intoken-report.ymlexplains the split clearly.concurrency: token-report-${{ inputs.org || 'petry-projects' }}withcancel-in-progress: trueis appropriate for a weekly cron.- The 65 000-char
MAXtruncation guard with a link-back to the run summary handles oversized reports defensively. - No shell injection surface: all interpolated values (
$repo,$id,$created_at) come fromgh apiJSON and are wrapped inprintf '%s'/ safe quoting.date -u -d "$created_at"has a|| echo 0fallback. - Find-or-create issue logic correctly paginates
listForRepoand matches on exact title.
Minor observations (non-blocking)
collect_org_jsonlusesgh api "orgs/${ORG}/repos?per_page=100&type=all" --paginate. For very large orgs this is O(repos) sequential artifact-list calls, as the docs already note — fine for current scale.- The
_extract_zipPython fallback is dead code on GitHub-hosted runners; harmless and cheap to keep.
CI status
All 27 check runs are SUCCESS or SKIPPED:
- AgentShield ✓
- CodeQL (actions) ✓
- SonarCloud — Quality Gate passed (0 new issues, 0 hotspots) ✓
- Lint: shellcheck (now covers
token_report.sh), bats (now runstests/token_report.bats), validate-agent-profiles, gh-aw-compile ✓ - CI: Lint, ShellCheck, Compile agentic workflows, Secret scan (gitleaks), Agent Security Scan ✓
- Tests: unit-tests ✓
- Dependency audit ✓
- Dev-Lead Agent dispatch ✓
- CodeRabbit and SonarCloud Code Analysis ✓
CodeRabbit hit a per-org rate limit and didn't post a substantive review, but its status is SUCCESS and the other static-analysis layers (CodeQL, SonarCloud, ShellCheck, gitleaks, AgentShield) all cleared.
mergeStateStatus: BLOCKED is solely due to required review, not failing checks.
Reviewed automatically by the PR-review agent (single-reviewer mode: opus 4.7). Reply if you need a human review.
There was a problem hiding this comment.
Pull request overview
Adds an org-wide “Token Cost Observatory” weekly report to make token-usage artifacts visible and actionable across all non-archived repos in the org (instead of being limited to .github-private), and reuses the same reporting logic in the daily fleet monitor Step Summary.
Changes:
- Introduces
scripts/token_report.shto collecttoken-usage-*artifacts org-wide and render aggregated Markdown (by workflow/tier/model and by repo). - Adds a scheduled workflow (
.github/workflows/token-report.yml) to post the weekly report as a comment on a tracking issue labeledtoken-report. - Updates the fleet monitor workflow to reuse the shared report script and extends CI to shellcheck + test it.
Reviewed changes
Copilot reviewed 9 out of 9 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
scripts/token_report.sh |
New org-wide artifact collector + Markdown report renderer used by both weekly report and fleet-monitor summary. |
tests/token_report.bats |
Unit tests for the pure aggregation/rendering functions (no network I/O). |
tests/fixtures/token_jsonl/run-a.jsonl |
JSONL fixture data for aggregation/rendering tests. |
tests/fixtures/token_jsonl/run-b.jsonl |
Additional JSONL fixture data for aggregation/rendering tests. |
.github/workflows/token-report.yml |
Weekly cron + manual dispatch workflow to generate and post the report to a tracking issue. |
.github/workflows/actions-fleet-monitor.yml |
Replaces single-repo token summary step with org-wide report via scripts/token_report.sh. |
.github/workflows/lint.yml |
Adds shellcheck coverage for the new script and runs the new bats tests. |
docs/token-report.md |
Documentation for the weekly report, ET metric, and operational usage. |
docs/actions-fleet-monitor.md |
Documents that fleet-monitor now includes an org-wide token usage rollup. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b5215c648b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Dev-Lead — fix-reviews (applied)Changes committed and pushed. |
e69063f
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e69063f731
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Dev-Lead — review-changes (applied)Changes committed and pushed. |
Dev-Lead — review-changes (applied)Changes committed and pushed. |
Dev-Lead — rate-limited (intent: fix-reviews)PR: #456 |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b96f7b8fa8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
Actionable comments posted: 4
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/lint.yml:
- Line 25: Update the shellcheck invocation under the run step that currently
reads "shellcheck --severity=warning scripts/fleet_monitor.sh
scripts/fleet_report.sh scripts/token_report.sh" to include explicit Bash mode
by adding the flag "--shell=bash" so the command becomes "shellcheck
--severity=warning --shell=bash ..." ensuring linting uses Bash semantics for
scripts/fleet_monitor.sh, scripts/fleet_report.sh, and scripts/token_report.sh.
In @.github/workflows/token-report.yml:
- Around line 51-57: The checkout step using
actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 with token: ${{
secrets.GH_PAT_WORKFLOWS }} persists the PAT in the repo's git config; update
the checkout "with" block for that step (the Checkout agent repo step) to set
persist-credentials: false so the PAT is not stored in the local git credential
helper after checkout.
- Around line 34-37: Remove the workflow-wide "issues: write" permission from
the top-level permissions block and instead grant "issues: write" only to the
specific job named "report" by adding a permissions block under jobs.report
(keep top-level permissions as minimal as required, e.g., contents: read) so
only the report job has write access to issues.
In `@scripts/token_report.sh`:
- Around line 233-235: The artifact download/extraction loop currently swallows
errors with bare `continue`, causing silent undercounts; update the failure
handlers for the `gh api "repos/${repo}/actions/artifacts/${id}/zip"` download
and the `_extract_zip "$zip" "$ex"` extraction to emit descriptive warnings to
stderr that include the repo, artifact id, and target paths (e.g., zip and ex)
and describe which step failed, then continue; reference the `gh api`
invocation, the `_extract_zip` call, and variables `repo`, `id`, `zip`, and `ex`
to locate where to add the warning messages.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 1266723d-8f5e-4314-876a-1da1272cf9d3
📒 Files selected for processing (9)
.github/workflows/actions-fleet-monitor.yml.github/workflows/lint.yml.github/workflows/token-report.ymldocs/actions-fleet-monitor.mddocs/token-report.mdscripts/token_report.shtests/fixtures/token_jsonl/run-a.jsonltests/fixtures/token_jsonl/run-b.jsonltests/token_report.bats
|
@coderabbitai resolve |
Dev-Lead — fix-reviews (applied)Changes committed and pushed. |
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…aily-digest accumulation (#1203) (#1204) * fix(fleet-monitor): auto-close resolved fleet-tracker issues (#1203) The monitor creates/updates a fleet-tracker issue per (repo × workflow) above 10% failure, but never closed them when the workflow recovered or was deleted — so 107 stale alerts accumulated (some with 'Last updated' stamps ~55 days old while the monitor runs daily; 4 for the deleted claude.yml #456). Add a step that closes any open fleet-tracker issue whose 'Last updated on <date>' stamp is older than STALE_DAYS (default 3 = 3 missed daily runs), with a ✅ Auto-resolved comment. It reopens automatically (fresh issue) if the workflow crosses the threshold again — same lifecycle as org-scorecard. Safety: gated on fleet_high_failure.json existing (a failed scan can't mass-close); staleness uses each issue's own stamp, not the current run, so a single bad-scan day never false-closes; comment + close are each guarded so a comment-capped issue still closes. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Juznz5V6su81ffSND8fg7s * fix(reviews): address review comments [skip ci-relay] * fix(fleet-monitor): close prior daily digests instead of accumulating them The 'workflow failures detected <date>' digest embeds the date in its title, so issues.create ran unconditionally every day — 42 open digests had piled up. Close the prior open digests (health-check label, matching title prefix) before opening today's, so only the current snapshot stays open. Complements the fleet-tracker auto-close in this PR; both stop the monitor accumulating stale issues. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Juznz5V6su81ffSND8fg7s * chore: dev-lead update (review-changes) [skip ci-relay] * fix(bot): address bot feedback [skip ci-relay] --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
…y) (#460) * fix(token-metrics): key usage sidecar per-call, not by $$ (concurrency) review-one-pr.sh backgrounds `run_agentic &` and `(run_duck) &` at the same time. $$ stays the parent PID inside both subshells, so the `${TOKEN_LOG_FILE}.last-usage.$$` sidecar collided between the two concurrent tier-2 calls — one engine's usage could overwrite or be read by the other before _record_engine_tokens logged it. (BASHPID doesn't work either: it differs between the chain's pipe subshell and the reader.) Fix: each run_* exports _ENGINE_USAGE_OUT, derived from its own per-call mktemp path (unique by construction) and inherited by the engine's pipeline subshell, so writer and reader agree while concurrent calls never share a file. $$ remains a fallback for non-concurrent direct callers. - token-metrics.sh: _engine_usage_sidecar prefers _ENGINE_USAGE_OUT. - engine.sh: run_triage/run_agentic/run_duck/run_writer export the per-call key. - tests: concurrency-isolation test (two jobs sharing $$ stay separate), per-call-key and fallback assertions; correct the misleading "$$ isolates parallel" test. - token_report.bats: make the malformed-JSONL test deterministic (jq 1.7 exits 0 on NUL bytes; use truncated JSON, which jq rejects across versions). Addresses PR #456 review (chatgpt-codex-connector P2: key usage sidecars by BASHPID). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * test(token-metrics): cover mktemp-failure fallback for per-call usage key Adds regression coverage for the fix that clears _ENGINE_USAGE_OUT before mktemp (PR #460 review, copilot-pull-request-reviewer ×3): - unit: a set-but-empty per-call key falls back to the $$-keyed sidecar (the mktemp-failure state), so it never reuses a prior/inherited key. - engine: a stale exported _ENGINE_USAGE_OUT + a forced mktemp failure does NOT reuse the stale sidecar — run_triage logs an estimate, not the planted 999/9/9/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
Why
We logged discussion #332 → issue #333 about token metrics, which shipped per-call JSONL logging (PR #334) and a summary step (PR #343). But the report was never actually visible:
REPO: petry-projects/.github-private. The agents run as reusable workflows in each caller repo, so theirtoken-usage-*artifacts land there. The summary saw 160 artifacts and was blind to ~2,300 others (markets, google-app-scripts, bmad-bgreat-suite, …).What
scripts/token_report.sh— discovers all non-archived repos, downloads everytoken-usage-*artifact in the lookback window, and renders an Effective Tokens (ET) rollup by workflow/tier/model and by repository. Purerender_*/aggregate_*functions (unit-tested);main()does the network I/O..github/workflows/token-report.yml— weekly cron (Mon 08:00 UTC) +workflow_dispatch. Posts the report as a comment on a single pinned tracking issue (labeltoken-report).actions-fleet-monitor.yml— the inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug; net −82 lines).tests/token_report.bats+ fixtures), docs (docs/token-report.md).Sample output (live, 2-day window)
The opus-4.7 pr-review tier dominates (15× multiplier × 4× output weight) — exactly the optimization signal the observatory was meant to surface.
Test plan
bats tests/token_report.bats(11 tests) +tests/fleet_report.batsshellcheck --severity=warningon all three scriptsCloses #206.
🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Documentation
Tests