Repository navigation
feat: token cost observatory — ET metric + JSONL logging - #334
Conversation
…w & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
Warning Rate limit exceeded
You’ve run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (9)
📝 WalkthroughWalkthroughThis PR instruments token-usage logging for pr-review and dev-lead agents. It adds a token-metrics library with ET calculation, modifies engine tier functions to capture output and record metrics, adds comprehensive unit tests, and configures GitHub Actions workflows to enable logging and upload JSONL artifacts. ChangesToken Metrics Instrumentation
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related issues
Possibly related PRs
🚥 Pre-merge checks | ✅ 5 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
CI StatusAll relevant checks pass ✅:
Pre-existing failures (not introduced by this PR):
@coderabbitai review |
There was a problem hiding this comment.
Pull request overview
Adds an opt-in “Token Cost Observatory” to capture per-LLM-call token estimates and compute Effective Tokens (ET), then persists usage logs as workflow artifacts for pr-review and dev-lead runs.
Changes:
- Introduces
scripts/lib/token-metrics.shfor model multipliers, ET calculation, token estimation, and JSONL emission. - Instruments
scripts/engine.shrun paths to estimate tokens (prompt/output) and emit one JSONL record per call whenTOKEN_LOG_FILEis set. - Wires workflow env + artifact upload to retain token usage logs for 30 days, and expands dev-lead test workflow path triggers.
Reviewed changes
Copilot reviewed 7 out of 7 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
tests/dev-lead/unit/test_token_metrics.bats |
Adds unit tests for multipliers, ET math, token estimation, and JSONL emission behavior. |
tests/dev-lead/unit/test_engine_writer.bats |
Adds integration assertions that writer/triage calls emit token records (and dry-run does not). |
scripts/lib/token-metrics.sh |
Implements ET + multipliers + estimate function + JSONL record emitter. |
scripts/engine.sh |
Adds token-record helper and tees model output to temp files for estimation-based logging. |
.github/workflows/test-dev-lead.yml |
Expands path filters to include scripts/lib/** changes. |
.github/workflows/pr-review.yml |
Enables token logging via env and uploads JSONL as an artifact. |
.github/workflows/dev-lead-reusable.yml |
Enables token logging via env and uploads JSONL as an artifact. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
🧩 Analysis chainI'll review the changes now. Let me read the relevant files first. 🏁 Script executed: Length of output: 57805 🏁 Script executed: Length of output: 22360
Token Cost Observatory — ReviewThe overall design is solid: opt-in, non-fatal, estimation-based, and well-tested. A few findings below. 🟡 Medium —
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8854ac840f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited - Add test 25: run_writer writes no token record when rate-limited (rc=2) Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing. 159/159 tests pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Actionable comments posted: 8
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/dev-lead-reusable.yml:
- Line 327: Replace the mutable tag "actions/upload-artifact@v4" with a pinned
full commit SHA (e.g., "actions/upload-artifact@<full-commit-sha>") so the
workflow uses an immutable reference; locate the repository’s existing pinned
SHA for actions/upload-artifact (used elsewhere in the repo) and copy that exact
commit SHA into the uses line to match the pinned-SHA approach already in the
project.
In @.github/workflows/pr-review.yml:
- Line 263: The workflow currently references reusable actions without digest
pins (uses: "actions/upload-artifact@v4" and "actions/cache@v5.0.5"); update
these uses to pinned commit digests by replacing "actions/upload-artifact@v4"
and "actions/cache@v5.0.5" with the corresponding full SHA commit DIDs (e.g.,
actions/upload-artifact@sha256:... and actions/cache@sha256:...) for the exact
release commits, verify the SHAs from the respective GitHub action
repositories/tags, and commit those digest-pinned strings back into
.github/workflows/pr-review.yml to ensure immutability of
"actions/upload-artifact" and "actions/cache".
In `@scripts/engine.sh`:
- Around line 350-351: The call to _record_engine_tokens in run_agentic is
hardcoding the tier as "deep" which mislabels audit/action runs; update the call
site to pass the actual agentic tier variable used by run_agentic (e.g. replace
the literal "deep" with the local tier variable such as "$AGENTIC_TIER" or
"$tier"), and if that variable isn't in scope, add it to the function parameters
or propagate it into scope so _record_engine_tokens receives the real tier;
ensure you only change the first argument to the actual tier and leave the other
arguments (e.g. "$REVIEW_ENGINE", "$model", "$prompt_file", "$_tok_tmp")
unchanged.
- Around line 337-338: The pipeline invocation using copilot_chat should be
changed to capture the copilot exit code without allowing set -e to abort before
rc is assigned: replace the two occurrences where you run copilot_chat
"$prompt_file" "$DEEP_TIMEOUT_SEC" --yolo | tee "$OUTPUT_FILE" followed by
rc=${PIPESTATUS[0]} with the pattern that appends "|| rc=${PIPESTATUS[0]}" to
the pipeline so the shell records the copilot_chat exit status reliably; update
both instances (the blocks invoking copilot_chat with OUTPUT_FILE at the two
places mentioned) and ensure the unique symbols copilot_chat, OUTPUT_FILE and
the use of PIPESTATUS[0] are preserved.
In `@scripts/lib/token-metrics.sh`:
- Around line 1-2: Add POSIX strict mode to the script by inserting "set -euo
pipefail" immediately after the shebang in token-metrics.sh so the script exits
on errors, treats unset variables as errors, and fails pipelines on the first
failing command; ensure the setting appears at the top of the file (right after
"#!/usr/bin/env bash") and does not alter other logic in the file.
- Around line 70-73: The JSONL record is built by interpolating raw shell
variables into record (the printf constructing '{"ts":...,"context":"%s"}'),
which breaks if context/model/workflow contain quotes or newlines; before
assembling record, escape/JSON-encode all string fields (at least context,
model, workflow, engine, tier, run_id) using a safe JSON-quoting helper (e.g.,
call jq -Rs ., python -c 'import json,sys;print(json.dumps(sys.argv[1]))', or a
dedicated shell function) and then use the escaped variables when building the
printf for record so the produced JSONL is always valid.
In `@tests/dev-lead/unit/test_engine_writer.bats`:
- Around line 27-28: The teardown currently unconditionally removes
$TOKEN_LOG_FILE which can point at caller-owned paths; instead ensure the test
only deletes files it created by (a) creating a test-owned temp file (e.g.
assign TOKEN_LOG_FILE="$(mktemp)" in the test setup) and (b) changing the
teardown to remove the file only if it matches that test-owned value or is
safely located in the test temp dir (e.g. check TOKEN_LOG_FILE is non-empty and
starts with "$TMPDIR/" and does not contain ".." before running rm -f).
Reference: TOKEN_LOG_FILE and the teardown snippet that currently runs rm -f
"$TOKEN_LOG_FILE".
In `@tests/dev-lead/unit/test_token_metrics.bats`:
- Around line 216-221: Add a regression test that calls emit_token_record with a
context string containing a double-quote and a newline (e.g. 'Line1
"quoted"\nLine2') and then read TOKEN_LOG_FILE and extract the context via jq -r
'.context' to assert it equals the original raw string; ensure you properly
quote/escape the input when calling emit_token_record in the test so the JSONL
output remains valid and jq returns the exact string (use bash $'...\\n...'
quoting or explicit escaping for the double-quote).
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: f57afb2a-2c09-4192-9cce-d55033c05897
📒 Files selected for processing (7)
.github/workflows/dev-lead-reusable.yml.github/workflows/pr-review.yml.github/workflows/test-dev-lead.ymlscripts/engine.shscripts/lib/token-metrics.shtests/dev-lead/unit/test_engine_writer.batstests/dev-lead/unit/test_token_metrics.bats
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a977979386
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- escape token JSONL emission with jq - add run_agentic tier parameter and call-site tiers - harden mktemp handling in engine paths - move/pin upload-artifact steps in workflows - tighten test teardown ownership and add JSON escaping regression - fix issue-triage workflow YAML run blocks Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
@coderabbitai resolve |
Dev-Lead — human-pr (applied)Changes committed and pushed. |
✅ Actions performedComments resolved and changes approved. |
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat(token-report): org-wide weekly Token Cost Observatory report The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call token-usage JSONL logging into the pr-review and dev-lead agents, but the only report was the fleet-monitor Step Summary — and it scanned only .github-private. Because the agents run as reusable workflows in each *caller* repo, their token-usage artifacts land in those repos, so the summary saw ~6% of real org spend and was buried where nobody looked. This adds org-wide collection and a weekly delivered report: - scripts/token_report.sh — discovers all non-archived repos, downloads every token-usage artifact in the lookback window, and renders an ET rollup by workflow/tier/model and by repository. Pure render_* functions are unit-tested; main() does the network I/O. - .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the report as a comment on a single pinned tracking issue (label: token-report). - actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug). - tests/token_report.bats + fixtures, docs/token-report.md, lint wiring. Closes #206. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * chore: apply manual instructions [skip ci-relay] * fix(reviews): address review comments [skip ci-relay] * feat(token-report): add effective-dated USD cost + unify ET with price table - scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows (price changes = append a dated row; calls priced at the rate on their own date). - scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date). - token-metrics.sh: model_multiplier_for now derives from the table (fixes stale opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart. - token_report.sh: annotate each record with date-accurate cost+ET; report now shows USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup; unpriced models surfaced (never silent $0). - tests: model_pricing.bats (incl. effective-date selection) + updated token_report and token_metrics expectations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(token-report): avoid SIGPIPE abort when trimming PR list The cost-per-PR section limited rows with `sort | head -10`. Under set -euo pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making the command substitution fail and aborting render_token_report — so no report is written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so sort never gets SIGPIPE. Adds a >10-PR test. Addresses PR #456 review (chatgpt-codex-connector P2). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(token-metrics): capture real token usage incl. cache (all engines) Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read was invisible everywhere. Now engine.sh captures real API usage when logging is on. - engine.sh: claude (chain + duck) and gemini run with --output-format json when TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and the real input / cache-read / cache-write / output counts are recorded. Gated by ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback to raw output if extraction is empty, so a parse hiccup never breaks a review. Usage crosses the `cmd | tee` subshell via a sidecar file. - copilot: gh copilot exposes no usage → stays on estimate (documented). - token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage; emit_token_record gains cache_creation_tokens (9th arg, default 0). - model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the report now price input + cache-read + cache-write + output. - stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reviews): address review comments [skip ci-relay] --------- Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead
- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)
Closes #333
Related: #332
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: verify fallback engine token capture when primary is rate-limited
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)
Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token observability review findings
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix token log path context in workflows
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: apply manual instructions [skip ci-relay]
* fix upload-artifact action pin
Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>



Summary
Implements the Token Cost Observatory described in Discussion #332 and planned in Issue #333.
What's new
scripts/lib/token-metrics.sh— core library: ET formula, model multipliers (haiku=1×, sonnet=3×, opus=15×, gemini-flash=0.5×), byte-based token estimation, JSONL emitterscripts/engine.sh— all 4 LLM tiers instrumented (run_triage,run_agentic,run_duck,run_writer) using tee-to-tmp capture +_record_engine_tokenshelper.github/workflows/pr-review.yml—TOKEN_LOG_FILE/TOKEN_WORKFLOW=pr-reviewenv vars + artifact upload (30-day retention).github/workflows/dev-lead-reusable.yml— same env vars + artifact upload forTOKEN_WORKFLOW=dev-lead.github/workflows/test-dev-lead.yml—scripts/lib/**added to path triggers--severity=warning -x)Design highlights
TOKEN_LOG_FILEis setts,workflow,tier,engine,model,input_tokens,cache_read_tokens,output_tokens,et,run_id,contextCloses #333
Related: #332
Summary by CodeRabbit
Release Notes
New Features
Tests