Skip to content

feat: token cost observatory — ET metric + JSONL logging - #334

Merged
don-petry merged 10 commits into
mainfrom
feat/token-cost-observatory
May 21, 2026
Merged

don-petry merged 10 commits into
mainfrom
feat/token-cost-observatory

Conversation

@don-petry

@don-petry don-petry commented May 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Implements the Token Cost Observatory described in Discussion #332 and planned in Issue #333.

What's new

  • scripts/lib/token-metrics.sh — core library: ET formula, model multipliers (haiku=1×, sonnet=3×, opus=15×, gemini-flash=0.5×), byte-based token estimation, JSONL emitter
  • scripts/engine.sh — all 4 LLM tiers instrumented (run_triage, run_agentic, run_duck, run_writer) using tee-to-tmp capture + _record_engine_tokens helper
  • .github/workflows/pr-review.yml — TOKEN_LOG_FILE / TOKEN_WORKFLOW=pr-review env vars + artifact upload (30-day retention)
  • .github/workflows/dev-lead-reusable.yml — same env vars + artifact upload for TOKEN_WORKFLOW=dev-lead
  • .github/workflows/test-dev-lead.yml — scripts/lib/** added to path triggers
  • 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
  • ShellCheck clean (--severity=warning -x)

Design highlights

  • Opt-in: zero overhead unless TOKEN_LOG_FILE is set
  • Non-fatal: token logging never aborts a workflow
  • Estimation-based (bytes ÷ 4): no dependency on API response JSON format
  • JSONL schema: ts, workflow, tier, engine, model, input_tokens, cache_read_tokens, output_tokens, et, run_id, context

Closes #333
Related: #332

Summary by CodeRabbit

Release Notes

  • New Features

    • Token usage metrics are now automatically collected and logged across development workflows for improved monitoring and cost tracking.
    • Generated token usage logs are retained as build artifacts for 30 days, enabling historical analysis and reporting.
  • Tests

    • Added comprehensive test coverage for token logging to ensure accurate metrics collection across different execution scenarios and edge cases.

Review Change Stack

…w & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings May 21, 2026 01:10
@gemini-code-assist

Copy link
Copy Markdown

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@coderabbitai

coderabbitai Bot commented May 21, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@don-petry has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 40 minutes and 38 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 381e0a7a-87b5-4345-ad2f-32a344335d8a

📥 Commits

Reviewing files that changed from the base of the PR and between a977979 and 05d249a.

📒 Files selected for processing (9)
  • .github/workflows/dev-lead-reusable.yml
  • .github/workflows/issue-triage-runner.yml
  • .github/workflows/pr-review.yml
  • docs/dev-lead/spec.md
  • scripts/engine.sh
  • scripts/lib/token-metrics.sh
  • scripts/review-one-pr.sh
  • tests/dev-lead/unit/test_engine_writer.bats
  • tests/dev-lead/unit/test_token_metrics.bats
📝 Walkthrough

Walkthrough

This PR instruments token-usage logging for pr-review and dev-lead agents. It adds a token-metrics library with ET calculation, modifies engine tier functions to capture output and record metrics, adds comprehensive unit tests, and configures GitHub Actions workflows to enable logging and upload JSONL artifacts.

Changes

Token Metrics Instrumentation

Layer / File(s) Summary
Token metrics library and unit tests
scripts/lib/token-metrics.sh, tests/dev-lead/unit/test_token_metrics.bats
New library defines model_multiplier_for, calculate_et, estimate_tokens_from_file, and emit_token_record functions. Tests validate model-to-multiplier mapping, ET formula correctness across input/cache/output combinations, ceiling-division token estimation, and JSONL record structure via jq validation. Edge-case tests confirm no-op behavior when TOKEN_LOG_FILE is unset and non-fatal failure handling for unwritable paths.
Engine token logging setup and helper
scripts/engine.sh (lines 12–17, 99–105, 192–209)
Engine documentation describes opt-in token logging via TOKEN_LOG_FILE and TOKEN_WORKFLOW. Library load is unconditional and non-fatal. New _record_engine_tokens helper conditionally emits JSONL records when logging is enabled and metrics functions are available, computing ET and including PR context.
Tier functions with token logging
scripts/engine.sh (run_triage, run_agentic, run_duck, run_writer)
Each tier function captures model output to a temp file via tee when token logging is enabled, calls _record_engine_tokens on success, and reliably cleans up temp files. Token recording is transparent to callers and does not change external behavior or parameters.
Engine token logging tests
tests/dev-lead/unit/test_engine_writer.bats (lines 27–28, 260–418)
Teardown extends TOKEN_LOG_FILE cleanup. Tests verify JSONL records are written with valid JSON and expected fields (tier, engine, model) for run_writer and run_triage. Tests confirm logging is skipped when disabled, dry-run produces no record, token capture to unwritable paths is non-fatal, fallback engine is correctly recorded, and rate-limited primary engine writes no token record.
Workflow environment and artifact upload
.github/workflows/dev-lead-reusable.yml (lines 54–55, 324–332), .github/workflows/pr-review.yml (lines 166–167, 260–268)
Both workflows set TOKEN_LOG_FILE (runner temp path keyed by run ID) and TOKEN_WORKFLOW environment variables. Each adds an unconditional actions/upload-artifact step that uploads the JSONL token log, ignores missing files, and retains artifacts for 30 days.
Workflow trigger path updates
.github/workflows/test-dev-lead.yml (lines 7, 17)
Added scripts/lib/** to both pull_request and push trigger path filters to ensure tests run when the token-metrics library changes.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related issues

  • Monitor and Optimize Token Usage #191: Shares the same token-usage instrumentation objectives (ET metric + JSONL logging) and workflow artifact management, so linking this PR would clarify the implementation of that discussion.

Possibly related PRs

  • petry-projects/.github-private#282: Modifies run_writer implementation to add Claude allowed tools; the token-logging instrumentation in this PR wraps the same writer execution path.
  • petry-projects/.github-private#227: Refactors run_writer_with_fallback and run_writer execution; this PR adds token-logging capture around the same functions.
  • petry-projects/.github-private#102: Modifies run_triage, run_agentic, and run_duck tier paths for Gemini and rate-limit handling; this PR instruments the same tier functions for token logging.
🚥 Pre-merge checks | ✅ 5 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 57.14% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'feat: token cost observatory — ET metric + JSONL logging' directly and specifically describes the main changes: introducing token cost tracking (observatory), the ET metric formula, and JSONL-based logging across workflows.
Linked Issues check ✅ Passed The PR implements all primary objectives from issue #333: token-metrics.sh library with ET formula and model multipliers, engine.sh instrumentation for all 4 LLM tiers, workflow TOKEN_LOG_FILE setup with artifact uploads, and comprehensive bats test coverage (28 library + 10 engine tests).
Out of Scope Changes check ✅ Passed All changes are directly aligned with issue #333 objectives. The workflow path-filter addition (scripts/lib/**) and test cleanup in test_engine_writer.bats are necessary supporting changes for the token metrics feature.
Description check ✅ Passed The PR provides comprehensive documentation through the commit message and PR description, including implementation details, test results, and acknowledgment of CodeRabbit's medium-severity findings regarding tier mapping in run_agentic and JSON escaping in emit_token_record.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/token-cost-observatory

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@don-petry

Copy link
Copy Markdown
Collaborator Author

CI Status

All relevant checks pass ✅:

  • ShellCheck ✓
  • bats unit tests (157/157) ✓
  • gh-aw-compile ✓
  • Secret scan ✓
  • CodeQL ✓
  • SonarCloud ✓
  • AgentShield ✓

Pre-existing failures (not introduced by this PR):

  1. CI/Lint — issue-triage-runner.yml line 30 YAML lint error (mapping values are not allowed here) — this failure exists on main before this branch was created. The :all: in the pip install command is unquoted YAML.

  2. Dev-Lead/dispatch — Claude rate-limited + gemini-1.5-pro ModelNotFoundError (deprecated model). This is an infrastructure/API issue unrelated to this PR.

@coderabbitai review

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds an opt-in “Token Cost Observatory” to capture per-LLM-call token estimates and compute Effective Tokens (ET), then persists usage logs as workflow artifacts for pr-review and dev-lead runs.

Changes:

  • Introduces scripts/lib/token-metrics.sh for model multipliers, ET calculation, token estimation, and JSONL emission.
  • Instruments scripts/engine.sh run paths to estimate tokens (prompt/output) and emit one JSONL record per call when TOKEN_LOG_FILE is set.
  • Wires workflow env + artifact upload to retain token usage logs for 30 days, and expands dev-lead test workflow path triggers.

Reviewed changes

Copilot reviewed 7 out of 7 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
tests/dev-lead/unit/test_token_metrics.bats Adds unit tests for multipliers, ET math, token estimation, and JSONL emission behavior.
tests/dev-lead/unit/test_engine_writer.bats Adds integration assertions that writer/triage calls emit token records (and dry-run does not).
scripts/lib/token-metrics.sh Implements ET + multipliers + estimate function + JSONL record emitter.
scripts/engine.sh Adds token-record helper and tees model output to temp files for estimation-based logging.
.github/workflows/test-dev-lead.yml Expands path filters to include scripts/lib/** changes.
.github/workflows/pr-review.yml Enables token logging via env and uploads JSONL as an artifact.
.github/workflows/dev-lead-reusable.yml Enables token logging via env and uploads JSONL as an artifact.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread scripts/engine.sh
Comment thread scripts/lib/token-metrics.sh Outdated
Comment thread .github/workflows/pr-review.yml Outdated
Comment thread .github/workflows/dev-lead-reusable.yml Outdated
@coderabbitai

coderabbitai Bot commented May 21, 2026

Copy link
Copy Markdown
🧩 Analysis chain

I'll review the changes now. Let me read the relevant files first.


🏁 Script executed:

#!/bin/bash
# Read the new token-metrics library
cat scripts/lib/token-metrics.sh

echo "=== ENGINE.SH (full) ==="
cat scripts/engine.sh

echo "=== PR-REVIEW WORKFLOW ==="
cat .github/workflows/pr-review.yml

echo "=== DEV-LEAD REUSABLE WORKFLOW ==="
cat .github/workflows/dev-lead-reusable.yml

echo "=== TEST-DEV-LEAD WORKFLOW ==="
cat .github/workflows/test-dev-lead.yml

echo "=== BATS TESTS FOR TOKEN-METRICS ==="
fd -e bats . --exec grep -l "token" {}

echo "=== LIST ALL BATS FILES ==="
fd -e bats .

Length of output: 57805


🏁 Script executed:

#!/bin/bash
echo "=== TOKEN METRICS TESTS ==="
cat tests/dev-lead/unit/test_token_metrics.bats

echo "=== ENGINE WRITER TESTS ==="
cat tests/dev-lead/unit/test_engine_writer.bats

echo "=== ENGINE FALLBACK TESTS ==="
cat tests/dev-lead/unit/test_engine_fallback.bats

echo "=== BATCH FALLBACK TESTS (for token integration) ==="
cat tests/test_batch_fallback.bats

Length of output: 22360


@don-petry: Great CI summary — all the pre-existing failures are clearly pre-existing and unrelated. Here's my review of the implementation.


Token Cost Observatory — Review

The overall design is solid: opt-in, non-fatal, estimation-based, and well-tested. A few findings below.


🟡 Medium — run_agentic hardcodes tier "deep", obscuring audit costs

_record_engine_tokens is called with the literal string "deep" in run_agentic, but the same function is used for both the deep-review and audit tiers (callers pass ENGINE_DEEP_MODEL vs ENGINE_AUDIT_MODEL). Audit tier is the most expensive (opus 4.7, multiplier 15×), so conflating it with "deep" makes per-tier cost breakdowns from the JSONL unreliable.

In engine.sh, the callers of run_agentic should forward the tier name as a parameter:

# run_agentic <prompt_file> <model> [tier]
run_agentic() {
  local prompt_file="$1"
  local model="$2"
  local tier="${3:-deep}"   # caller passes "deep", "audit", or "action"
  ...
  if [ "$rc" -eq 0 ]; then
    _record_engine_tokens "$tier" "$REVIEW_ENGINE" "$model" "$prompt_file" "$_tok_tmp"
  fi
  ...
}

🟡 Medium — JSON injection risk in emit_token_record

emit_token_record builds the JSONL record using printf '%s' with no JSON-escaping applied to any field:

record=$(printf \
  '{"ts":"%s","workflow":"%s","tier":"%s","engine":"%s","model":"%s",...,"context":"%s"}' \
  "$ts" "$workflow" "$tier" "$engine" "$model" \
  "$input" "$cache" "$output" "$et" "$run_id" "$context")

The context field is ${PR_URL:-}. While standard GitHub PR URLs are safe, branch names in query strings or non-default PR_URL values could contain " or \, silently producing malformed JSON that fails jq empty on the collected log. The existing test only covers a clean URL.

Consider building the record with jq -n:

record=$(jq -n \
  --arg ts "$ts" \
  --arg workflow "$workflow" \
  --arg tier "$tier" \
  --arg engine "$engine" \
  --arg model "$model" \
  --argjson input_tokens "$input" \
  --argjson cache_read_tokens "$cache" \
  --argjson output_tokens "$output" \
  --argjson et "$et" \
  --arg run_id "$run_id" \
  --arg context "$context" \
  '{ts:$ts,workflow:$workflow,tier:$tier,engine:$engine,model:$model,
    input_tokens:$input_tokens,cache_read_tokens:$cache_read_tokens,
    output_tokens:$output_tokens,et:$et,run_id:$run_id,context:$context}')

This also keeps jq as the single JSON dependency (already present in all runners).


🔵 Minor — ci-relay artifact upload is dead code

In dev-lead-reusable.yml, TOKEN_LOG_FILE is set only in the dispatch job's env: block. The ci-relay job also has the upload step but TOKEN_LOG_FILE is never injected there, so the file never exists and the step is always a no-op (if-no-files-found: ignore). It won't cause failures, but it adds ~1s and confusion. Can be removed from ci-relay, or TOKEN_LOG_FILE can be added to that job's env if future relay-tier cost tracking is desired.


✅ Nit — Redundant exact match arm in model_multiplier_for

o4-mini | *o4-mini*)        echo "2.0" ;;

The first o4-mini exact-match arm is fully subsumed by the *o4-mini* fnmatch glob below it in the same | alternative. The glob would also match openai/o4-mini (the actual COPILOT_API_MODEL value). Can drop the exact arm, though it's entirely harmless.


✅ Positive highlights

  • Non-fatal design is rigorous: every exit path from _record_engine_tokens and emit_token_record is guarded, || true is applied correctly, and the unwritable-path test confirms it.
  • PIPESTATUS[0] usage is correct throughout: the tee pipeline correctly captures the LLM CLI's exit code, not tee's.
  • run_writer reuse of _tmp for both rate-limit detection and token estimation is efficient and correct — the asymmetry with other tiers' _tok_tmp is justified since _tmp serves a second purpose.
  • Test coverage is thorough: ET formula arithmetic, ceiling division edge cases (1 char, 3 chars), multi-record append, no-op when unset, non-fatal on unwritable path — all covered.
  • Opt-in zero-overhead guarantee is clean: declare -f emit_token_record >/dev/null guard in _record_engine_tokens ensures safety even if the lib fails to source.
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8854ac840f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread .github/workflows/dev-lead-reusable.yml Outdated
Comment thread scripts/engine.sh Outdated
Comment thread scripts/engine.sh Outdated
don-petry and others added 2 commits May 20, 2026 20:20
- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 8

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/dev-lead-reusable.yml:
- Line 327: Replace the mutable tag "actions/upload-artifact@v4" with a pinned
full commit SHA (e.g., "actions/upload-artifact@<full-commit-sha>") so the
workflow uses an immutable reference; locate the repository’s existing pinned
SHA for actions/upload-artifact (used elsewhere in the repo) and copy that exact
commit SHA into the uses line to match the pinned-SHA approach already in the
project.

In @.github/workflows/pr-review.yml:
- Line 263: The workflow currently references reusable actions without digest
pins (uses: "actions/upload-artifact@v4" and "actions/cache@v5.0.5"); update
these uses to pinned commit digests by replacing "actions/upload-artifact@v4"
and "actions/cache@v5.0.5" with the corresponding full SHA commit DIDs (e.g.,
actions/upload-artifact@sha256:... and actions/cache@sha256:...) for the exact
release commits, verify the SHAs from the respective GitHub action
repositories/tags, and commit those digest-pinned strings back into
.github/workflows/pr-review.yml to ensure immutability of
"actions/upload-artifact" and "actions/cache".

In `@scripts/engine.sh`:
- Around line 350-351: The call to _record_engine_tokens in run_agentic is
hardcoding the tier as "deep" which mislabels audit/action runs; update the call
site to pass the actual agentic tier variable used by run_agentic (e.g. replace
the literal "deep" with the local tier variable such as "$AGENTIC_TIER" or
"$tier"), and if that variable isn't in scope, add it to the function parameters
or propagate it into scope so _record_engine_tokens receives the real tier;
ensure you only change the first argument to the actual tier and leave the other
arguments (e.g. "$REVIEW_ENGINE", "$model", "$prompt_file", "$_tok_tmp")
unchanged.
- Around line 337-338: The pipeline invocation using copilot_chat should be
changed to capture the copilot exit code without allowing set -e to abort before
rc is assigned: replace the two occurrences where you run copilot_chat
"$prompt_file" "$DEEP_TIMEOUT_SEC" --yolo | tee "$OUTPUT_FILE" followed by
rc=${PIPESTATUS[0]} with the pattern that appends "|| rc=${PIPESTATUS[0]}" to
the pipeline so the shell records the copilot_chat exit status reliably; update
both instances (the blocks invoking copilot_chat with OUTPUT_FILE at the two
places mentioned) and ensure the unique symbols copilot_chat, OUTPUT_FILE and
the use of PIPESTATUS[0] are preserved.

In `@scripts/lib/token-metrics.sh`:
- Around line 1-2: Add POSIX strict mode to the script by inserting "set -euo
pipefail" immediately after the shebang in token-metrics.sh so the script exits
on errors, treats unset variables as errors, and fails pipelines on the first
failing command; ensure the setting appears at the top of the file (right after
"#!/usr/bin/env bash") and does not alter other logic in the file.
- Around line 70-73: The JSONL record is built by interpolating raw shell
variables into record (the printf constructing '{"ts":...,"context":"%s"}'),
which breaks if context/model/workflow contain quotes or newlines; before
assembling record, escape/JSON-encode all string fields (at least context,
model, workflow, engine, tier, run_id) using a safe JSON-quoting helper (e.g.,
call jq -Rs ., python -c 'import json,sys;print(json.dumps(sys.argv[1]))', or a
dedicated shell function) and then use the escaped variables when building the
printf for record so the produced JSONL is always valid.

In `@tests/dev-lead/unit/test_engine_writer.bats`:
- Around line 27-28: The teardown currently unconditionally removes
$TOKEN_LOG_FILE which can point at caller-owned paths; instead ensure the test
only deletes files it created by (a) creating a test-owned temp file (e.g.
assign TOKEN_LOG_FILE="$(mktemp)" in the test setup) and (b) changing the
teardown to remove the file only if it matches that test-owned value or is
safely located in the test temp dir (e.g. check TOKEN_LOG_FILE is non-empty and
starts with "$TMPDIR/" and does not contain ".." before running rm -f).
Reference: TOKEN_LOG_FILE and the teardown snippet that currently runs rm -f
"$TOKEN_LOG_FILE".

In `@tests/dev-lead/unit/test_token_metrics.bats`:
- Around line 216-221: Add a regression test that calls emit_token_record with a
context string containing a double-quote and a newline (e.g. 'Line1
"quoted"\nLine2') and then read TOKEN_LOG_FILE and extract the context via jq -r
'.context' to assert it equals the original raw string; ensure you properly
quote/escape the input when calling emit_token_record in the test so the JSONL
output remains valid and jq returns the exact string (use bash $'...\\n...'
quoting or explicit escaping for the double-quote).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: f57afb2a-2c09-4192-9cce-d55033c05897

📥 Commits

Reviewing files that changed from the base of the PR and between 8ec3f19 and a977979.

📒 Files selected for processing (7)
  • .github/workflows/dev-lead-reusable.yml
  • .github/workflows/pr-review.yml
  • .github/workflows/test-dev-lead.yml
  • scripts/engine.sh
  • scripts/lib/token-metrics.sh
  • tests/dev-lead/unit/test_engine_writer.bats
  • tests/dev-lead/unit/test_token_metrics.bats

Comment thread .github/workflows/dev-lead-reusable.yml Outdated
Comment thread .github/workflows/pr-review.yml Outdated
Comment thread scripts/engine.sh Outdated
Comment thread scripts/engine.sh Outdated
Comment thread scripts/lib/token-metrics.sh
Comment thread scripts/lib/token-metrics.sh Outdated
Comment thread tests/dev-lead/unit/test_engine_writer.bats Outdated
Comment thread tests/dev-lead/unit/test_token_metrics.bats

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a977979386

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/lib/token-metrics.sh Outdated
don-petry and others added 5 commits May 21, 2026 06:33
- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@don-petry

Copy link
Copy Markdown
Collaborator Author

@coderabbitai resolve

@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — human-pr (applied)

Changes committed and pushed.

coderabbitai[bot]
coderabbitai Bot previously approved these changes May 21, 2026
@coderabbitai

coderabbitai Bot commented May 21, 2026

Copy link
Copy Markdown
✅ Actions performed

Comments resolved and changes approved.

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@sonarqubecloud

Copy link
Copy Markdown

@don-petry
don-petry merged commit b5100b7 into main May 21, 2026
25 checks passed
@don-petry
don-petry deleted the feat/token-cost-observatory branch May 21, 2026 11:50
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Jun 25, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 2, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 2, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 3, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat(token-report): org-wide weekly Token Cost Observatory report

The Token Cost Observatory (discussion #332, PRs #334/#343) wired per-call
token-usage JSONL logging into the pr-review and dev-lead agents, but the only
report was the fleet-monitor Step Summary — and it scanned only .github-private.
Because the agents run as reusable workflows in each *caller* repo, their
token-usage artifacts land in those repos, so the summary saw ~6% of real org
spend and was buried where nobody looked.

This adds org-wide collection and a weekly delivered report:

- scripts/token_report.sh — discovers all non-archived repos, downloads every
  token-usage artifact in the lookback window, and renders an ET rollup by
  workflow/tier/model and by repository. Pure render_* functions are unit-tested;
  main() does the network I/O.
- .github/workflows/token-report.yml — weekly cron (Mon 08:00 UTC) that posts the
  report as a comment on a single pinned tracking issue (label: token-report).
- actions-fleet-monitor.yml — its inline single-repo summary now reuses the shared
  script, so the daily Step Summary is org-wide too (fixes the hardcoded-repo bug).
- tests/token_report.bats + fixtures, docs/token-report.md, lint wiring.

Closes #206.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* chore: apply manual instructions [skip ci-relay]

* fix(reviews): address review comments [skip ci-relay]

* feat(token-report): add effective-dated USD cost + unify ET with price table

- scripts/lib/model-pricing.tsv: single source of truth, effective-dated rows
  (price changes = append a dated row; calls priced at the rate on their own date).
- scripts/lib/model-pricing.sh: price_for / cost_usd / et_multiplier_for (glob+date).
- token-metrics.sh: model_multiplier_for now derives from the table (fixes stale
  opus=15 → 5; Opus 4.5+ is $5 input). ET and USD can no longer drift apart.
- token_report.sh: annotate each record with date-accurate cost+ET; report now shows
  USD cost by workflow/tier/model and by repo, plus a most-expensive-PRs rollup;
  unpriced models surfaced (never silent $0).
- tests: model_pricing.bats (incl. effective-date selection) + updated token_report
  and token_metrics expectations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(token-report): lint model-pricing.sh + model_pricing.bats; document cost layer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(token-report): avoid SIGPIPE abort when trimming PR list

The cost-per-PR section limited rows with `sort | head -10`. Under set -euo
pipefail, head closing the pipe early can leave sort with SIGPIPE (141), making
the command substitution fail and aborting render_token_report — so no report is
written or posted. Use `awk 'NR<=10'` instead: it consumes the full stream, so
sort never gets SIGPIPE. Adds a >10-PR test.

Addresses PR #456 review (chatgpt-codex-connector P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(token-metrics): capture real token usage incl. cache (all engines)

Token counts were estimated (char/4) and cache was hardcoded to 0, so cache-read
was invisible everywhere. Now engine.sh captures real API usage when logging is on.

- engine.sh: claude (chain + duck) and gemini run with --output-format json when
  TOKEN_LOG_FILE is set; the model's text is extracted for downstream consumers and
  the real input / cache-read / cache-write / output counts are recorded. Gated by
  ENGINE_USAGE_JSON (default on; set 0 to revert to text+estimate). Robust fallback
  to raw output if extraction is empty, so a parse hiccup never breaks a review.
  Usage crosses the `cmd | tee` subshell via a sidecar file.
- copilot: gh copilot exposes no usage → stays on estimate (documented).
- token-metrics.sh: parse_engine_usage / extract_engine_text / reset_engine_usage;
  emit_token_record gains cache_creation_tokens (9th arg, default 0).
- model-pricing: add cache_write column (5m write = 1.25x input); cost_usd + the
  report now price input + cache-read + cache-write + output.
- stubs gain a JSON usage mode; tests cover parsing, cache capture end-to-end, the
  ENGINE_USAGE_JSON kill-switch, gemini usage, and cache-write pricing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reviews): address review comments [skip ci-relay]

---------

Co-authored-by: donpetry-bot <{}+donpetry-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: donpetry-bot <281750570+donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 7, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
don-petry added a commit that referenced this pull request Aug 8, 2026
* feat: token cost observatory — ET metric + JSONL logging for pr-review & dev-lead

- Add scripts/lib/token-metrics.sh with ET formula, model multipliers, JSONL emitter
- Instrument engine.sh run_triage/run_agentic/run_duck/run_writer with token capture
- Always load token-metrics.sh at engine.sh source time (non-fatal; no-op if file absent)
- Add TOKEN_LOG_FILE + TOKEN_WORKFLOW env vars to pr-review.yml and dev-lead-reusable.yml
- Artifact upload (token-usage-${run_id}.jsonl, 30-day retention) to both workflows
- Add scripts/lib/** path triggers to test-dev-lead.yml
- 38 new bats unit tests (28 library + 10 engine integration); all 157 tests pass
- ShellCheck clean (--severity=warning -x)

Closes #333
Related: #332

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test: verify fallback engine token capture when primary is rate-limited

- Add test 24: run_writer_with_fallback logs gemini (not claude) when claude rate-limited
- Add test 25: run_writer writes no token record when rate-limited (rc=2)

Both cases verified: fallback success logs correct engine; rate-limit failure logs nothing.
159/159 tests pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token observability review findings

- escape token JSONL emission with jq
- add run_agentic tier parameter and call-site tiers
- harden mktemp handling in engine paths
- move/pin upload-artifact steps in workflows
- tighten test teardown ownership and add JSON escaping regression
- fix issue-triage workflow YAML run blocks

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix token log path context in workflows

Use /tmp token log paths in job env and artifact upload paths to satisfy actionlint context rules.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: apply manual instructions [skip ci-relay]

* fix upload-artifact action pin

Use the repository-standard pinned SHA for actions/upload-artifact (v7.0.1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: donpetry-bot <donpetry-bot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: Token Cost Observatory — ET metric + JSONL logging for pr-review & dev-lead agents

3 participants