Skip to content

fix(claude-ops): give the skill-visibility starvation band a real usage signal - #3532

Merged
kyle-sexton merged 24 commits into
mainfrom
feat/usage-tracking-claude-json
Sep 1, 2026
Merged

kyle-sexton merged 24 commits into
mainfrom
feat/usage-tracking-claude-json

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Closes #3534

Summary

Started as an exploration of what Claude Code tracks in ~/.claude.json, and whether this repo was about to rebuild a counter the product already keeps. It found two defects in claude-ops:audit-skill-visibility, both fixed here, and established that the mechanism both this repo and the official docs describe is not the one the product implements.

Fix

The starvation band carried no usage signal. compute_listing sorted on a usage_score that only the test fixtures ever set. The live collector builds its denominator from a filesystem walk, and usage events were joined afterwards, so every real run scored zero and the band fell through to its tiebreaker while presenting itself as usage-informed. On this machine that ranked adhd:clarify (1 use) as first to lose its description and work-items:triage (99 uses) among the safest. Scores are now computed before the listing is built.

Usage recorded under a skill's bare leaf was silently discarded. Events were looked up by qualified <plugin>:<leaf> name only, while the stores hold both that key and the bare leaf as separate rows. source-control:babysit-prs reported 97 invocations against an actual 475. A bare key is now attributed when exactly one skill owns that leaf, and withheld with its candidates when more than one does, which is the same refusal the custom_skill redaction guard already makes.

The mechanism, recovered and then corrected twice. listing_score mirrors the product's own scorer, usageCount * max(0.5 ** (daysSinceUse / 7), 0.1). Two further corrections came out of an independent re-extraction by a fresh-context adjudicator:

  • Truncation is a greedy first-fit walk, not a score-ordered prefix. The grant loop has no early exit, so a cheap low-scored description can be granted after an expensive higher-scored one was refused. Description length is a second ranking input that no prose account of this mechanism mentions.
  • The grant budget is computed forward from a floor, budget - V, where V is what the listing costs before any description is granted. Without it the grant boundary itself is wrong, which would have left the first-fit fix inert.
  • Ties keep catalog order, not alphabetical. The product's sort is stable. This decides everything in the unscored case, where every score is zero and the tiebreak is the whole ordering.

listing.score_basis reports unscored when no usage survives to weigh, and competing rows carry confidence: "unscored" rather than borrowing inferential.

What is deliberately not changed: demand/overflow still ignores name bytes. It answers "does the listing overflow", a different question from "which entries win", and only the second needs the floor. At this repo's 12.7x overflow, charging names cannot flip the verdict, but it would move every test encoding the current arithmetic and every figure recorded against it. It wants its own evidence.

ADR 0016 gets two dated revision blockquotes in its own established shape. The core decision is untouched in both.

  • The Context line's drop order is corrected, and the error is attributed upstream: the sentence restates https://code.claude.com/docs/en/skills ("Skill descriptions are cut short"), which is itself wrong and still wrong as of 2026-08-31. That is why the same sentence keeps re-entering this repo; it was on main in the SKILL.md and the manifest, and docs: replace in-place skill-frontmatter restatements with upstream references #3524 added a citation pointing at the wrong page for it.
  • The deferral clause's ground is restated on three reasons documentation cannot cure (the scorer is wrong-signed for the question, skillUsage cannot name never-invoked skills, it carries no take-up attribution), because characterizing the substrate invites a false lift. Lift conditions recorded.

Documentation. Two new references beside the skill that consumes them: reference/listing-scorer.md (the mechanism, its counterexamples, the drift posture) and reference/usage-counters.md (what skillUsage, pluginUsage, agentLastUsed and projects[] actually mean, including the 60-second throttle that drops rather than coalesces, the pluginUsage install seeding, and the structural absence of per-project skill usage).

Verification

  • Python suite 94/94, ten new cases covering the scorer, the decay boundaries, the unscored basis, bare-key attribution, the ambiguous-key refusal, the first-fit walk, and the name floor.
  • scripts/run-ruff.sh check clean; markdownlint clean on every touched file.
  • Live re-run confirms the fixes: score_basis: native-counters, source-control:babysit-prs at 475, clean correctly withheld as ambiguous between disk-hygiene:clean and repo-hygiene:clean.
  • The verification stamp fired on the day it was written. The CLI auto-updated 2.1.251 to 2.1.252 mid-work; both the scorer and the truncator re-verified unchanged. The re-run surfaced a trap now recorded in all three stamps: minified identifiers are not stable across builds (zPe became WPe, and three sibling symbols moved with it), so a recheck must locate by shape, not by name.
  • skill-quality:check reports one error and one warning, both pre-existing: audit_skill_visibility.test.sh exits 2 before and after this change from an unrelated fixture gap, and the 1024-codepoint description warning predates it (the description is now shorter than it started).

Related

  • ADR 0016, docs/adr/0016-source-skill-recommendation-from-the-catalog-not-the-listing.md, amended here but not superseded.
  • docs/conventions/upstream-drift, the convention the two new verification stamps follow.
  • docs/conventions/topic-docs, whose contract-slice rule pruned this work's task doc; its durable outcomes graduated into the two references above, ADR 0016, and audit-skill-visibility: starvation band is alphabetical, and bare usage keys are dropped #3534.
  • docs/native-surfaces/records.json gains the doctor to audit-skill-visibility row (complementary), recorded through the generator and independently confirmed.

🤖 Generated with Claude Code

https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d

kyle-sexton and others added 12 commits August 31, 2026 11:48
Exploration artifact only. Records the verified shape and write paths of
skillUsage, pluginUsage, agentLastUsed and projects[] in ~/.claude.json, the
recovered skill-listing budget scorer, and the repo surfaces that already read
those counters.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
A live run of audit_skill_visibility.py returns usage_score 0 for every row, so
the band falls back to its alphabetical tiebreaker. Records the run's numbers
and the inverted example rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
The observability reporter's lanes omit ~/.claude.json; nothing reads
agentLastUsed or the per-project cost snapshot; pair co-occurrence needs the
event stream the counters cannot supply.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
Adds the Claim/Basis/As-of/Recheck stamp for the binary-derived listing scorer
and the fallback rule for a build mismatch. Restates the skill-usage.jsonl
finding as what was checked and what remains unresolved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…ints

Records claude_code.skill_activated's invocation_trigger as the only source
that separates user-invoked from model-invoked, the retro transcript parser as
the only working agent-usage source, and the three codified refusals: ADR 0016's
deferral of usage-driven surfacing, the performance engine's stat-only posture
on ~/.claude.json, and the exposure-floor withheld verdict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…plit drop

The audit is a three-source reconciler with a capability tier model, not a
native-counter reader, so parse_otel and TIER_CAPABILITIES replace the earlier
claim that OTEL appears only in a reference file. The 60s debounce was already
documented at SKILL.md:194-196, so it is recorded as not-a-gap rather than a
finding. The qualified-vs-bare key split is confirmed: classify() looks events
up by qualified_name alone, so bare-key rows are discarded silently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…no re-derivation

Section 3.9 was doing audit-native-overlap's job in prose, a silent second way
alongside docs/native-surfaces/records.json. It now raises the missing
doctor -> audit-skill-visibility row as a candidate and leaves the verdict to
the human gate that skill's contract requires.

Part 3.8 told the reader to start from the codified refusals "rather than
re-derive them", which is the inertia this repo's incumbency discipline exists
to catch. It now separates the two refusals that name a purpose and a
measurement from the one that names only the state of the substrate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…ad both ways

A cross-vendor re-derivation run blind to the earlier reasoning found the
"scope limit" framing wrong in both directions. The deferral is scoped to
show-options rotation, so encoding zPe inside audit-skill-visibility was never
inside it and does not need to be framed as an exception. But the deferral is
also not up for lifting: its real ground is that zPe is wrong-signed for a
forgotten-skill nudge, that skillUsage cannot name never-invoked skills, and
that it carries no take-up attribution. Records the lift conditions and the
reason a binary-derived stamp is not what the ADR meant by documented.

Also drafts, without applying, the correction owed to the ADR's Context line
about drop ordering.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…al ground

Two dated revision blockquotes in the ADR's own established shape, which
preserves superseded reasoning rather than editing Context in place.

The Context paragraph stated that Claude Code drops descriptions "starting with
the skills invoked least". The mechanism is a decay-weighted score,
usageCount * max(0.5 ^ (daysSinceUse / 7), 0.1), sorted descending and granted
greedily, so a heavily used but stale skill can be shed before a lightly used
fresh one. The correction strengthens the decision rather than weakening it: a
never-invoked skill still scores zero and is shed first, and the decay term adds
a second bias against exactly the forgotten-skill population show-options
exists to surface.

The deferral clause rested on skillUsage being "undocumented internal state".
That substrate is now characterized and dated, which invites a false lift, so
the ground is restated on three reasons documentation cannot cure: the scorer is
wrong-signed for the question, skillUsage cannot name the never-invoked
population, and it carries no take-up attribution. Lift conditions recorded.

Core decision untouched in both cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…pping bare usage keys

Two defects in audit-skill-visibility, both confirmed against live data.

compute_listing sorted on a usage_score only the test fixtures ever set. The
live collector builds its denominator from a filesystem walk and joined usage
events afterwards, so every real run scored zero and the band fell through to
its alphabetical tiebreaker while presenting itself as usage-informed. On this
machine that ranked adhd:clarify (1 use) first to lose its description and
work-items:triage (99 uses) among the safest. Scores are now computed before
the listing is built.

Events were looked up by qualified <plugin>:<leaf> name only, while the stores
hold both that key and the bare leaf as separate rows, so the bare row was
discarded silently. source-control:babysit-prs reported 97 invocations against
an actual 475. A bare key is now attributed when exactly one skill owns that
leaf and withheld with its candidates when more than one does, which is the
same refusal the custom_skill redaction guard makes.

listing_score mirrors the product's own scorer,
usageCount * max(0.5 ** (daysSinceUse / 7), 0.1), recovered from Claude Code
2.1.251 and carrying a verification stamp with its basis and recheck trigger.
It is fed from native counters under the exact qualified key because the
product's scorer does no bare-key fallback either; scoring the merged total
would predict a truncation that will not happen.

listing.score_basis reports "unscored" when no usage survives to weigh, and
competing rows carry confidence "unscored" rather than borrowing "inferential",
which would claim more than an alphabetical order supports.

Suite 93/93, ruff clean. The shell harness exits 2 both before and after this
change, from a pre-existing fixture gap unrelated to it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
The counting-rules additions pushed SKILL.md from 199 to 216 lines, past the
soft target, which is the progressive-disclosure signal working. The mechanism,
why "least invoked" was the wrong description, why a bare key does not move the
band, and the drift posture now live in reference/listing-scorer.md, leaving two
tight bullets and a pointer in the hub.

SKILL.md lands at 207. The remaining eight lines are the two non-obvious
counting rules the section exists to hold, so they stay rather than being cut to
hit an advisory number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
The rewritten drop-order clause pushed description+when_to_use to 1550 chars
against the 1536 cap, which truncates the listing entry. Says the same thing in
fewer words: descriptions are dropped by a decay-weighted usage score.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1e61d7713d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

kyle-sexton and others added 3 commits August 31, 2026 17:10
Matches the citation shape of the ADR's two 2026-08-21 revisions, which carry an
issue link. 3534 records the two audit-skill-visibility defects that prompted
both revisions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…mentary verdict

audit-skill-visibility's description already routed the one-shot
unused-versus-context-cost check to the bundled /doctor, with no store
row behind the disclaimer. Record the missing row: verdict
complementary, evidence from the installed v2.1.252 binary's doctor
Check 1 strings (2026-08-31), and regenerate NATIVE-SURFACES.md via
overlap.py generate. Fresh-context verifier concurred with the verdict;
self-check degraded only on the pre-existing v2.1.232 extraction
staleness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
Resolves two conflicts, both mechanical:

plugin.json: keep 0.39.0. This branch's minor bump supersedes main's 0.38.22
patch; the two changes are independent and both ship.

CHANGELOG.md: keep both sections, 0.39.0 above 0.38.22. Neither entry replaces
the other.

Suite 93/93 on the merged tree.

NOTE, unresolved by this merge and tracked separately: main's #3524 added a
SKILL.md paragraph and a plugin.json description clause restating the
least-invoked-first drop order and citing the official skills page for it. This
branch's own description now says decay-weighted, so the merged file states both.
An adjudicator is settling the wording; the contradiction is deliberate and
visible rather than silently resolved one way here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 57s —— View job


Security review — /review:security-review

  • Gather context (PR diff, changed files, existing Codex review comments)
  • Invoke /review:security-review skill and follow its procedure
  • Post findings

Scope: logic flaws, authorization/access-control gaps, injection surfaces, unsafe handling of tokens/secrets/credentials, and dangerous GitHub Actions patterns, per the skill's charter. Style, correctness, and general quality are out of scope for this lane (that's /review:code-review; Codex's three P2 comments on this PR are correctness/quality, not security, and are left for that lane).

Result: no security issues found.

Reviewed the full diff (docs/NATIVE-SURFACES.md, ADR 0016, docs/native-surfaces/records.json, the new docs/topics/usage-tracking-claude-json/EXPLORE.md, plugins/claude-ops/.claude-plugin/plugin.json, CHANGELOG.md, SKILL.md, the new reference/listing-scorer.md, and scripts/audit_skill_visibility.py / its test file):

  • No GitHub Actions workflow files are touched by this PR — nothing to assess for pull_request_target/workflow_run risk, permissions: widening, or script injection via the github context.
  • scripts/audit_skill_visibility.py's new/changed code (listing_score, resolve_event_keys, compute_listing's scoring wiring, classify's native-score computation) is pure, read-only data reconciliation over already-parsed usage events (~/.claude.json skillUsage, the plugin's own JSONL store, OTEL). No new file paths, subprocess calls, eval/exec, or shell interpolation are introduced. The pre-existing subprocess.run call in the file (around line 691) is untouched by this diff and outside its changed lines.
  • The bare-key attribution logic (resolve_event_keys) only withholds or attributes counts within the already-collected fleet; it can't be used to inject or misattribute data across a trust boundary — worst case per the Codex comments is a report-accuracy edge case (duplicate qualified names across marketplaces), which is a correctness concern, not a security one.
  • No secrets, tokens, or credentials are read, logged, or newly exposed. The new EXPLORE.md documents shapes of ~/.claude.json fields (skillUsage, pluginUsage, etc.) and the governance section explicitly reaffirms the existing safety rule that ~/.claude.json and history.jsonl stay stat-only/allowlisted elsewhere in the audit engine; this PR doesn't change what fields audit_skill_visibility.py reads (still only firstStartTime and skillUsage).
  • Nothing here changes authorization, trust boundaries, or introduces network calls; this remains a local, read-only diagnostic CLI a user runs against their own machine state.

No inline comments posted — there is nothing to anchor in this lane.

@github-actions

github-actions Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Last security-reviewed head: c96caec418bfd043f4963b8667f49dd88617f021. On the next push, the relevance gate compares only the commits since this SHA; delete this comment to force a full re-review.

@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 39s —— View job


Claude finished @kyle-sexton's task

Ran /review:code-review (org CI code-review lane) against origin/main...HEAD at 460be55, scoped to correctness/maintainability per this lane's split (security is out of scope here).

  • Gather context (PR diff, CLAUDE.md, review skill)
  • Review changed files for correctness/maintainability
  • Cross-check the three existing Codex findings
  • Post inline comments for confirmed findings
  • Final summary

Findings

1. score_basis should be computed from competing rows only — confirmed, posted inline.
Independently verified Codex's finding at audit_skill_visibility.py:830. score_basis is decided over all rows, but native_scores in classify iterates the whole denominator with no eligibility filter, so a bundled/name-only/disable-model-invocation skill (excluded from competing) with real native usage flips the basis to native-counters even when every actual competing row scores zero — mislabeling an alphabetical order as inferential instead of unscored. That's the exact class of bug this PR's summary says it eliminated, reintroduced one layer up.

2. The default markdown render never reads score_basis — confirmed by direct read, could not anchor inline (untouched line, outside the diff hunk).
audit_skill_visibility.py:1258-1275. The --render markdown path (the report's default surface) unconditionally states "descriptions are dropped least-invoked-first" and frames the ordering as an inferential "likelihood band," even when listing["score_basis"] == "unscored" — the exact "alphabetical order presenting itself as usage-informed" problem the PR fixes in the JSON model, left standing in the default rendered output. The wording is also stale independent of that: "least-invoked-first" is the mechanism this PR corrects everywhere else, and "undocumented scorer" is no longer accurate now that listing_score is documented.

3. Bare-key attribution for duplicate qualified names — reviewed Codex's finding, not elevating.
audit_skill_visibility.py:158-179, resolve_event_keys. Codex is right that leaf_owners[leaf] is a set, so two denominator rows sharing one qualified name (the two-marketplaces case) dedupe to a single owner, and a bare-key event then attributes to both rows via starvation_by_name/events_by_skill.get(name). But this reproduces the existing, tested behavior for duplicate qualified names under direct qualified-key events too (AmbiguousAttributionTest, test_duplicate_qualified_name_is_marked_ambiguous): both rows already receive identical counts and are already flagged attribution: "ambiguous-attribution" on output. The bare-key path is consistent with that accepted design rather than a new class of error, so I'd call this a real edge case worth a code comment but not a blocking defect.

Not re-litigated

Everything else in the diff (the ADR revision blockquotes, EXPLORE.md, listing-scorer.md, resolve_event_keys's own-key/ambiguity logic apart from finding 3, the new test coverage) checked out on read — no further correctness issues found within this lane's scope.

@github-actions

github-actions Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude has reviewed this PR 5 times. The lane skips further automatic reviews after 5; deleting this comment resets the count.

…ocate-by-shape

The recheck trigger fired the same day it was written: the CLI auto-updated from
2.1.251 to 2.1.252 mid-session. Re-ran the extraction. Both the scorer's formula
and the truncator's descending-sort greedy grant come back unchanged.

The re-run surfaced a trap worth recording. The minified identifier is not stable
across builds: the scorer was zPe in 2.1.251 and WPe in 2.1.252 with a
byte-identical body. A recheck that greps the old name finds nothing and would
wrongly conclude the mechanism was removed. All three stamps now say to locate by
shape, and carry the two greps that do it.

Suite 93/93, ruff clean, markdownlint clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 41s —— View job


Security review — /review:security-review

  • Gather context (PR diff since last reviewed head, changed files)
  • Invoke /review:security-review skill and follow its procedure
  • Post findings

Result: skip gate — no new security-relevant surface since the last review.

The last successful security review on this PR covered head 460be55 (see comment above). The only commit since then is 0aaf24e ("re-verify the listing scorer at 2.1.252 and stamp locate-by-shape"):

git diff --stat 460be55..0aaf24e
 docs/topics/usage-tracking-claude-json/EXPLORE.md            |  9 ++--
 .../reference/listing-scorer.md                              | 33 +++++++++-------
 .../scripts/audit_skill_visibility.py                        | 14 ++++---
 3 files changed, 40 insertions(+), 16 deletions(-)

I read the audit_skill_visibility.py hunk directly (it's the only non-.md file touched): the change is confined to a comment block above listing_score — updating the reverse-engineering provenance note to record re-verification against Claude Code 2.1.252 and add a "locate by shape, not by minified name" grep tip. No executable statements changed. The other two files are documentation only.

This matches skip-gate condition 3 in /review:security-review's procedure ("this head already has a successful security review that still applies"), so per the skill's instructions I'm stopping here and posting nothing further. The full-diff review already posted above (no security issues found: no Actions workflows touched, no new file paths/subprocess/eval/shell interpolation, no secrets exposure, read-only local diagnostic) still stands for 0aaf24e18656c5b0b8f198efecbedeb35cb34e2b.

@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 24s —— View job


Code review — /review:code-review

  • Gather context (PR diff, changed files, prior Codex/Claude review comments)
  • Invoke /review:code-review skill and follow its procedure
  • Check whether the prior review from this lane still applies
  • Final summary

Result: skip gate hit — no new findings to post.

This run is triggered by the synchronize that landed HEAD at 0aaf24e. I diffed 460be55 (the commit this lane already reviewed at 22:30 UTC) against 0aaf24e and it touches only three files, all comments/docs:

  • audit_skill_visibility.py — verification-stamp comment block only (re-verifies the scorer at Claude Code 2.1.252, adds a locate-by-shape note). No code changed.
  • reference/listing-scorer.md and docs/topics/usage-tracking-claude-json/EXPLORE.md — same re-verification, in prose.

Per the skill's skip gate ("this head already has a successful review from this lane that still applies"), that review still applies. I re-verified its two confirmed findings by reading the current file directly, since neither line moved:

  1. score_basis computed from all rows, not competing — still present at audit_skill_visibility.py:836 (score_basis = "native-counters" if any(r["usage_score"] for r in rows) else "unscored", evaluated before competing is filtered at line 838). Already flagged inline on this PR; not re-posting a duplicate.
  2. Default markdown render never branches on score_basis — still present at audit_skill_visibility.py:1257-1281: still unconditionally says "descriptions are dropped least-invoked-first" and "an undocumented scorer," even when listing["score_basis"] == "unscored", and "undocumented" is now stale on top of that since listing_score is documented in this PR's own reference/listing-scorer.md.

No new correctness/maintainability issues found in the delta itself. Both open findings above are unaddressed as of this HEAD and worth resolving before merge, but since they're unchanged from the prior review, no new inline comments are posted here to avoid duplicating what's already on the PR.

…opic slice

Two gates caught real violations of this repo's own conventions.

contract-slice-prune-gate: docs/topics/<slug>/ is Contract tier per
docs/conventions/topic-docs/README.md, committed on a task branch and pruned
before merge. EXPLORE.md was the wrong home from the start; that convention also
names EXPLORE.md as a memory-tier file, which never gets committed at all. The
durable outcomes are graduated rather than deleted:

- The scorer, its drift posture and the locate-by-shape rule already live in
  reference/listing-scorer.md.
- The counters' own semantics now live in a new reference/usage-counters.md,
  beside the skill that consumes them: the skillUsage throttle and its
  qualified-versus-bare key split, the pluginUsage install seeding that makes
  lastUsedAt worthless at zero count, the incomparability of pluginUsage across
  plugin shapes, agentLastUsed holding nothing usable, and the structural fact
  that no per-project skill usage exists in this file at all.
- The two defects are issue 3534; the mechanism correction is ADR 0016.

Every citation of the pruned slice is repointed, so nothing dangles.

skill-count-claim-gate: a bullet opening "One skill, two possible usage keys"
parsed as a claim that claude-ops ships one skill. Reworded to "Two possible
usage keys per skill", same meaning, no leading count word.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 8s —— View job


Code review — /review:code-review

  • Gather context (PR diff, CLAUDE.md, review skill)
  • Invoke /review:code-review skill and follow its procedure
  • Check whether the prior review from this lane still applies
  • Post findings

Result: skip gate hit (condition 4) — no new findings to post.

This run is triggered by the synchronize that landed HEAD at 9930754 (full SHA 993075405a4bad137c09b841b40e6e8c53ad23ec). I diffed 0aaf24e (the commit this lane already reviewed at 22:50 UTC) against HEAD:

git diff --stat 0aaf24e..HEAD
 docs/adr/0016-...-not-the-listing.md               |   9 +-
 docs/topics/usage-tracking-claude-json/EXPLORE.md  | 546 ---------------------
 plugins/claude-ops/CHANGELOG.md                    |   4 +-
 .../skills/audit-skill-visibility/SKILL.md         |   2 +-
 .../reference/listing-scorer.md                    |   5 +-
 .../reference/usage-counters.md                    | 130 +++++
 .../scripts/audit_skill_visibility.py              |   4 +-

All seven files are documentation/comment moves: EXPLORE.md's content is retired in favor of the new reference/usage-counters.md (added in the prior head, at 0aaf24e, and unchanged since), and the ADR, CHANGELOG, SKILL.md, and listing-scorer.md edits only repoint cross-references from the deleted EXPLORE.md path to the new reference/ files. The one hunk in audit_skill_visibility.py is confined to the same provenance comment block reviewed last time — it now points at reference/listing-scorer.md/reference/usage-counters.md instead of EXPLORE.md. No executable statement changed.

Per the skill's skip gate ("this head already has a successful review from this lane that still applies"), that review still applies. Re-verified its two confirmed findings directly against HEAD, since neither moved:

  1. score_basis computed from all rows, not competing — still present at audit_skill_visibility.py:838 (score_basis = "native-counters" if any(r["usage_score"] for r in rows) else "unscored", evaluated before competing is filtered). Already flagged inline on this PR by both Codex and this lane; not re-posting a duplicate.
  2. Default markdown render never branches on score_basis — still present at audit_skill_visibility.py:1259-1283: still unconditionally says "descriptions are dropped least-invoked-first" and attributes the ordering to "an undocumented scorer," even when listing["score_basis"] == "unscored" — and "undocumented" is stale on top of that, since listing_score is now documented in this PR's own reference/listing-scorer.md.

No new correctness/maintainability issues found in the delta itself. Both open findings above remain unaddressed as of this HEAD and worth resolving before merge, but since they're unchanged from the prior review, no new inline comments are posted here to avoid duplicating what's already on the PR.

@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 2s —— View job


Security review — /review:security-review

  • Gather context (PR diff since last reviewed head, changed files)
  • Invoke /review:security-review skill and follow its procedure
  • Post findings

Result: skip gate — no new security-relevant surface since the last review.

The last successful security review on this PR covered head 0aaf24e18656c5b0b8f198efecbedeb35cb34e2b. The only commit since then is 9930754 ("graduate the usage-counter findings and prune the topic slice"):

git diff --stat 0aaf24e..9930754
 .../0016-source-skill-recommendation-from-the-catalog-not-the-listing.md | 9 +-
 docs/topics/usage-tracking-claude-json/EXPLORE.md                        | 546 -----
 plugins/claude-ops/CHANGELOG.md                                          | 4 +-
 plugins/claude-ops/skills/audit-skill-visibility/SKILL.md                | 2 +-
 .../reference/listing-scorer.md                                         | 5 +-
 .../reference/usage-counters.md                                         | 130 +++
 .../scripts/audit_skill_visibility.py                                   | 4 +-
 7 files changed, 144 insertions(+), 556 deletions(-)

audit_skill_visibility.py is again the only non-.md file touched, and the diff is confined to a comment block above listing_score: it retargets provenance pointers from the now-removed docs/topics/usage-tracking-claude-json/EXPLORE.md (deleted in this commit) to the new reference/listing-scorer.md and reference/usage-counters.md files. No executable statements changed.

The remaining changes are documentation reorganization: EXPLORE.md's content graduated into the new reference/usage-counters.md, the ADR gets a cross-reference update, and CHANGELOG.md/SKILL.md get matching pointer updates. I read the new reference/usage-counters.md directly — it documents the shapes of ~/.claude.json fields (skillUsage, pluginUsage, agentLastUsed, projects[]) and reaffirms the same stat-only/allowlisted governance already covered in the prior review; it introduces no new code paths, commands, or data flows.

This matches skip-gate condition 3 in /review:security-review's procedure ("this head already has a successful security review that still applies"). The full-diff review posted earlier on this PR (no security issues found: no Actions workflows touched, no new file paths/subprocess/eval/shell interpolation, no secrets exposure, read-only local diagnostic) still stands for the current head 993075405a4bad137c09b841b40e6e8c53ad23ec.

…re prefix

The adjudicator's independent re-extraction found the mirror wrong on the half I
had not checked. The product's grant loop has NO early exit:

  for (let me of W) {
    let ge = me.entryLen - (me.cmd.name.length + 2);
    if (ge <= pe) pe -= ge; else fe.push(me);
  }

So truncation is a greedy first-fit walk over the whole score-descending list,
not a prefix of it. A cheap low-scored description can be granted after an
expensive higher-scored one was refused, which makes description LENGTH a second
ranking input that no prose account of this mechanism mentions. compute_listing
modelled a prefix, which understated the exposure of long descriptions and
overstated it for short ones.

The verdict now mirrors the real walk. The band keeps ranking exposure by score,
and the two are allowed to disagree; a regression test pins exactly that case,
where a band-1 row survives a pass that sheds a better-scored longer row.

Budget accounting is deliberately NOT changed here. It counts description bytes
against the whole budget and ignores the name bytes every entry also pays. That
is the CERTAIN half of the report and a separate correction with its own
evidence.

Also corrects the attribution. The ADR's drop-order sentence tracks the official
documentation, which states the same false claim verbatim ("drops descriptions
starting with the skills you invoke least, so the skills you use most keep their
full text"). The revision now names upstream as the source of the error instead
of implying this repo got it wrong, and the SKILL.md paragraph that cited that
page as authority now says which part of it does not hold. The plugin.json
description clause is corrected the same way.

Suite 94/94, ruff clean, markdownlint clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 48s —— View job


Code review — /review:code-review

  • Gather context (PR diff, changed files, prior review comments)
  • Invoke /review:code-review skill and follow its procedure
  • Post inline findings
  • Final summary

Ran the code-review lane against HEAD 916367b (org CI code-review split: correctness/maintainability; security is a separate lane and already cleared this PR). This run's trigger is the synchronize that landed 916367b ("model the listing truncation as first-fit, not a score prefix") on top of 9930754, the last commit this lane reviewed. I focused on that delta.

What changed since the last review

compute_listing's verdict/band logic in audit_skill_visibility.py was rewritten: the prior model shed the lowest-scored prefix of competing skills until overflow was covered; the new model walks every competing entry in descending-score order with a running budget and no early exit — the actual first-fit mechanism the product runs, per the accompanying reference/listing-scorer.md and ADR 0016 updates. I traced the logic by hand (audit_skill_visibility.py:844-868) against the new test_a_cheap_low_scored_row_is_granted_after_a_costly_higher_one_is_shed test and it's correct — band (exposure ranking, ascending) and verdict (first-fit walk, descending) are now properly decoupled, and the test's own arithmetic checks out. Good fix.

New findings (posted inline)

  1. verify-posttooluse-probe.md — a stray scratch file at the repo root (deliberate "teh" typo, trailing whitespace, unreferenced anywhere else in the repo). Reads as a leftover PostToolUse hook probe accidentally committed; suggest dropping it.
  2. plugins/claude-ops/CHANGELOG.md — the [0.39.0] entry documents the score-timing and bare-key fixes but not this commit's first-fit-vs-prefix correction, which changes real verdict output for real runs. Worth its own bullet given how thorough the rest of the entry is.

Still open from prior rounds (unchanged by this commit, not re-posted to avoid duplicating)

Neither touched by 916367b; already flagged inline earlier in this PR by both Codex and this lane:

  • score_basis is decided over all rows, not just competing — audit_skill_visibility.py:840, can mislabel an alphabetical order as inferential instead of unscored.
  • The default --render markdown path never branches on score_basis — audit_skill_visibility.py:1259-1283 — still states "least-invoked-first" and an "undocumented scorer" unconditionally, which is now stale prose too.

Not run

Couldn't execute the Python test suite (scripts/affected-tests.sh --run / pytest) in this session — Bash execution beyond git requires interactive approval that wasn't available here. Verified the new logic by static trace against the new test case instead; if you want the suite actually executed, that needs re-running with broader tool permissions.
· branch feat/usage-tracking-claude-json

@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 1s —— View job


Security review — /review:security-review

  • Gather context (PR diff since last reviewed head, changed files)
  • Invoke /review:security-review skill and follow its procedure
  • Post findings

Result: skip gate — no new security-relevant surface since the last review.

The last successful security review on this PR covered head 993075405a4bad137c09b841b40e6e8c53ad23ec. The only commit since then is 916367b ("model the listing truncation as first-fit, not a score prefix"):

git diff --stat 9930754..916367b
 .../0016-source-skill-recommendation-from-the-catalog-not-the-listing.md | 24 ++++++-----
 plugins/claude-ops/.claude-plugin/plugin.json                            |  2 +-
 plugins/claude-ops/skills/audit-skill-visibility/SKILL.md                | 19 ++++----
 .../reference/listing-scorer.md                                         | 22 +++++++++-
 .../scripts/audit_skill_visibility.py                                   | 61 +++++++++++++-------
 .../scripts/test_audit_skill_visibility.py                              | 49 +++++++++++++++++
 verify-posttooluse-probe.md                                              |  4 ++
 7 files changed, 138 insertions(+), 43 deletions(-)

I read every hunk directly:

  • audit_skill_visibility.py — reworks compute_listing's truncation model from a score-ordered prefix to a greedy first-fit walk (verdict) plus a separate exposure ranking (band), and updates the verification-stamp comment to match. Pure in-memory arithmetic over already-collected local data; no new file paths, subprocess calls, network access, or trust-boundary changes. budget (line 814) is defined before its use at line 860, so no latent bug either.
  • test_audit_skill_visibility.py — one new unit test exercising the first-fit behavior. No security surface.
  • ADR 0016, SKILL.md, reference/listing-scorer.md, plugin.json description string — documentation/prose updates describing the corrected mechanism. No code.
  • verify-posttooluse-probe.md (new, 4 lines) — a static markdown file with two lines of placeholder text ("teh" typos, trailing whitespace), unrelated to the stated PR purpose. It executes nothing and isn't referenced by any script or hook in this diff, so it carries no security implication; it reads like leftover local test debris and may be worth dropping before merge, but that's a housekeeping note for the code-review lane, not a security finding.

This matches skip-gate condition 3 in /review:security-review's procedure ("this head already has a successful security review that still applies"). The full-diff review posted earlier on this PR (no security issues found: no Actions workflows touched, no new file paths/subprocess/eval/shell interpolation, no secrets exposure, read-only local diagnostic) still stands for the current head 916367b5bec5a88ace2beefb4d4431af94f26582.

verify-posttooluse-probe.md is a live probe fixture belonging to the in-flight
skill-usage-hook investigation, deliberately containing a misspelling and
trailing whitespace to trip a PostToolUse hook. It has no business in history.
The file stays on disk so the running worker is not disrupted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
Comment thread verify-posttooluse-probe.md Outdated
Comment thread plugins/claude-ops/CHANGELOG.md
@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 20s —— View job


Code review — /review:code-review

  • Gather context (PR diff, CLAUDE.md, review skill)
  • Invoke /review:code-review skill and follow its procedure
  • Check whether the prior review from this lane still applies
  • Post findings
  • Final summary

This run is triggered by the synchronize that landed HEAD at 0c63eb8. Diffed against 9930754 (the commit this lane last reviewed): two new commits, 916367b (a real logic rewrite — models the listing truncation as greedy first-fit instead of a score-ordered prefix) and 0c63eb8 (untracks a scratch probe file, unrelated). This is not a skip-gate cycle — 916367b touches compute_listing's core algorithm — so I did a full review rather than re-affirming the prior one.

New in this cycle: the first-fit rewrite (916367b) — looks correct

compute_listing (audit_skill_visibility.py:842-889) now sorts competing descending by score and walks every row against a running budget with no early exit (likely-retained if it fits in what's left, likely-starved otherwise), separately from the band exposure ranking (ascending by score). This matches the recovered JS walk documented in reference/listing-scorer.md, and the new test (test_a_cheap_low_scored_row_is_granted_after_a_costly_higher_one_is_shed) exercises the case that motivates the change — a cheap row surviving on leftover budget after a costlier, better-scored row is shed. I traced the arithmetic by hand against that test and it checks out. SKILL.md, the ADR, and listing-scorer.md were updated in the same commit to describe first-fit instead of a prefix, consistently. No issues found in this part of the diff.

Findings

1. score_basis computed from all rows, not competing — still open, posted inline.
Unchanged by this cycle. Previously flagged by Codex and this lane on 460be55/0aaf24e/9930754. At audit_skill_visibility.py:840, score_basis is decided over rows before competing is filtered at line 842, and classify()'s native_scores walks the whole denominator with no eligibility filter — so a bundled/name-only/disable-model-invocation skill with real native usage can flip score_basis to "native-counters" while every actual competing row scores zero, mislabeling an alphabetical order as inferential instead of unscored.

2. Default markdown render never reads score_basis, and is now the last stale corner of this skill — could not anchor inline (untouched by this PR, outside its diff hunks).
audit_skill_visibility.py:1274-1298. The --render markdown path — the report's default surface — still unconditionally says descriptions are "dropped least-invoked-first" (line 1286) and that the order "comes from an undocumented scorer" (line 1295), never branching on listing["score_basis"]. This was already open on prior heads; 916367b makes it more visibly stale, not less: that same commit rewrote SKILL.md, the ADR, and reference/listing-scorer.md to say "sheds descriptions lowest-scoring first" and cites the now-documented scorer, but left this render function — in the same file, ~700 lines below its own updated verification-stamp comment — using the exact phrasing the rest of the PR corrects.

3. Bare-key attribution for duplicate qualified names — reviewed Codex's finding again, not elevating (unchanged from prior review).
resolve_event_keys still collapses two denominator rows sharing one qualified name (the two-marketplaces case) to a single owner via a set, so a bare-key event attributes to both. This reproduces the existing, tested behavior for duplicate qualified names under direct qualified-key events too (both rows already get identical counts and attribution: "ambiguous-attribution"), so it's consistent with the accepted design rather than a new defect class.

Not re-litigated

The rest of the diff since 9930754 — plugin.json's description sync, the ADR/SKILL.md/listing-scorer.md prose updates, and the new test — checked out on read.
· branch feat/usage-tracking-claude-json

…, and date it

The correction stays at full strength, but "stated wrongly" with no referent put
the error on the ADR's author, who transcribed the vendor page near-verbatim. It
also dropped the operationally useful fact: the upstream page is still wrong as
of 2026-08-31, which is why the same sentence keeps re-entering this repo. It is
on main in the skill's SKILL.md and in the plugin manifest, and #3524 added a
citation pointing at the wrong page for it.

The SKILL.md paragraph now separates the two upstream pages by whether they hold:
the settings page owns the budget fraction and per-entry cap and matches; the
skills page states the drop order and does not. Readers are routed to the binary
for the ordering.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
@claude

claude Bot commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 48s —— View job


Security review — /review:security-review

  • Gather context (PR diff since last reviewed head, changed files)
  • Invoke /review:security-review skill and follow its procedure
  • Post findings

Result: skip gate — no new security-relevant surface since the last review.

The last successful security review on this PR covered head 916367b5bec5a88ace2beefb4d4431af94f26582. Two commits landed since then, bringing HEAD to c8b80ff:

git diff --stat 916367b..c8b80ff
 .../0016-source-skill-recommendation-from-the-catalog-not-the-listing.md | 12 ++++++------
 plugins/claude-ops/skills/audit-skill-visibility/SKILL.md                | 13 +++++++------
 verify-posttooluse-probe.md                                              |  4 ----
 3 files changed, 13 insertions(+), 16 deletions(-)

I read both hunks directly:

  • verify-posttooluse-probe.md deleted (commit 0c63eb8) — this is the stray scratch file flagged in this lane's own prior review comment; it's now removed from the tree entirely. No security implication either way (it was inert markdown), but confirms the housekeeping note was actioned.
  • ADR 0016 and SKILL.md (commit c8b80ff) — prose-only edits reattributing the drop-order documentation error to upstream (code.claude.com/docs/en/skills) more precisely and dating the "still wrong as of 2026-08-31" claim. No code, no new links to untrusted content, no change to what the audit script reads or executes.

No executable code changed in this delta. This matches skip-gate condition 3 in /review:security-review's procedure ("this head already has a successful security review that still applies"). The full-diff review posted earlier on this PR (no security issues found: no Actions workflows touched, no new file paths/subprocess/eval/shell interpolation, no secrets exposure, read-only local diagnostic) still stands for the current head c8b80ff2f76dcccc552c086600865c7abc97bff3.

Two more corrections from the adjudicator's re-extraction, both scoped to the
grant loop.

The grant budget is computed FORWARD from a floor, `budget - V`, never backward
from the overflow. V is what the listing costs before any description is
granted: every listed entry pays for its own name, the exempt classes pay their
full rendering because they are never candidates, and the separators are charged
too. A disable-model-invocation skill is absent from the listing entirely, so
unlike the other two exempt classes it costs nothing and takes no separator.
Without this the grant boundary itself is wrong, which left the first-fit fix in
916367b inert.

Ties keep catalog order, not alphabetical. The product's sort is stable, so equal
scores stay in input order, and Python's is too. This decides everything in the
unscored case, where every score is zero and the tiebreak IS the whole ordering;
an alphabetical one disagreed with the product on every row.

demand/overflow is deliberately NOT changed. It answers "does the listing
overflow", a different question from "which entries win", and only the second
needs the floor. Charging name bytes there is real but not urgent: at this
repo's 12.7x overflow it cannot flip the verdict, and it would move every test
encoding the current arithmetic plus every figure recorded against it. It wants
its own evidence, a fleet where the corrected count changes the verdict.

Two test expectations move because the old ones encoded the no-floor model, and
the corrected answers are what the product actually does. Live run: 171 of 176
competing now starved, up from 167, which is the floor being charged.

Also softens listing-scorer.md's "a never-used skill is always shed first". It
sorts last, so it sheds first under any material overflow, but first-fit means a
zero-scored skill with a very short description can still be granted.

Suite 94/94, ruff clean, markdownlint clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
@claude

claude Bot commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 54s —— View job


Security review — /review:security-review

  • Gather context (PR diff since last reviewed head, changed files)
  • Invoke /review:security-review skill and follow its procedure
  • Post findings

Result: no security issues found.

The last successful security review on this PR covered head c8b80ff2f76dcccc552c086600865c7abc97bff3. One commit landed since then, bringing HEAD to ef9d226 (fix(claude-ops): charge the name floor and keep catalog order on ties):

git diff --stat c8b80ff..ef9d226
 .../reference/listing-scorer.md                    | 21 ++++++++---
 .../scripts/audit_skill_visibility.py              | 41 ++++++++++++++++------
 .../scripts/test_audit_skill_visibility.py         | 28 ++++++++++-----
 3 files changed, 68 insertions(+), 22 deletions(-)

This is a real code change (not skip-gate-eligible on docs alone), so I read the full delta directly:

  • audit_skill_visibility.py — two corrections inside compute_listing, both pure in-memory arithmetic over already-collected local data: (1) the grant budget (remaining) is now computed forward as budget - floor, where floor charges every listed entry for its own qualified-name length and separator bytes (plus full rendering cost for bundled/exempt rows), instead of being derived backward from overflow; (2) the competing/by_exposure sort keys drop the qualified_name tiebreaker, relying on Python's stable sort to preserve catalog (input) order for equal scores instead of forcing alphabetical order. Neither change touches file I/O, subprocess invocation, network access, or any trust boundary — it's recomputing rankings and budget bytes from data already parsed earlier in the script. No new inputs are introduced; entry["qualified_name"] and description lengths are used only for arithmetic and sort keys, never interpolated into a command, path, or template.
  • test_audit_skill_visibility.py — updated/renamed unit test assertions matching the new floor-and-catalog-order arithmetic. No security surface.
  • reference/listing-scorer.md — prose documenting the same two corrections. No code.

No GitHub Actions workflow files are touched by this commit or by the PR as a whole. This matches skip-gate condition 3 in spirit (no new security-relevant surface), and the full-diff review posted earlier on this PR (no security issues found: no Actions workflows touched, no new file paths/subprocess/eval/shell interpolation, no secrets exposure, read-only local diagnostic) still stands for the current head ef9d226e116c2166672b2c6ae13f41d53e01d90a.
· branch feat/usage-tracking-claude-json

kyle-sexton and others added 2 commits August 31, 2026 20:47
docs/CATALOG.md is generated from the plugin manifests, and the drop-order
wording fix in the claude-ops description left it stale. Regenerated with
scripts/generate-catalog.mjs; the only diff is that one clause.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
Refreshes a stale base the gate correctly refused. main had moved two commits
touching the same claude-ops manifest and changelog, which is exactly the
overlap that lets a stale-base squash silently revert a landed fix.

Both conflicts mechanical again:

plugin.json: keep 0.39.0 and this branch's corrected description clause. main
still carries "drops descriptions least-invoked-first", which is the upstream
wording this branch fixes, so taking theirs would have reverted the correction.

CHANGELOG.md: keep all three entries, 0.39.0 above main's new 0.38.23 above
0.38.22. None replaces another.

Catalog already in sync, suite 94/94 on the merged tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
@claude

claude Bot commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 3s —— View job


Security review — /review:security-review

  • Gather context (PR diff since last reviewed head, changed files)
  • Invoke /review:security-review skill and follow its procedure
  • Post findings

Result: skip gate — no new security-relevant surface since the last review.

The last successful security review on this PR covered head ef9d226e116c2166672b2c6ae13f41d53e01d90a. HEAD is now d65e747 via a merge commit:

git log --oneline ef9d226..d65e747
d65e747 Merge origin/main into feat/usage-tracking-claude-json
79f1c29 fix: keep skill argument substitution out of embedded shell text (#3541)
f7e51dc docs: regenerate the catalog after the claude-ops description change

79f1c29 and f7e51dc are commits that landed on origin/main independently and were pulled in by this merge — confirmed via git merge-base origin/main HEAD, which resolves to 79f1c29 itself. The PR's actual diff (git diff origin/main...HEAD) is unchanged in scope from prior reviews: the same 11 files (docs/CATALOG.md, docs/NATIVE-SURFACES.md, ADR 0016, docs/native-surfaces/records.json, plugins/claude-ops/.claude-plugin/plugin.json, plugins/claude-ops/CHANGELOG.md, SKILL.md, reference/listing-scorer.md, reference/usage-counters.md, audit_skill_visibility.py, test_audit_skill_visibility.py), 817/-42.

Diffing the PR-owned files between ef9d226 and d65e747 directly shows only two touched, both prose:

  • docs/CATALOG.md — one line, rewording the claude-ops catalog description's drop-order phrase ("least-invoked-first" → "lowest-score-first"), consistent with this PR's own correction.
  • plugins/claude-ops/CHANGELOG.md — one new entry ([0.38.23]) documenting an unrelated observability skill fix (a stray $1 placeholder in dispatch prose), which is main's own change merged in, not new work from this branch.

audit_skill_visibility.py and every other file under audit-skill-visibility/ are byte-identical to the last reviewed head (git diff --stat ef9d226..d65e747 -- 'plugins/claude-ops/skills/audit-skill-visibility/**' returns empty). No executable code changed, no GitHub Actions workflow files are touched, and no new file paths, subprocess/eval calls, secrets, or trust-boundary changes were introduced.

This matches skip-gate condition 3 in /review:security-review's procedure ("this head already has a successful security review that still applies"). The full-diff review posted earlier on this PR (no security issues found: no Actions workflows touched, no new file paths/subprocess/eval/shell interpolation, no secrets exposure, read-only local diagnostic) still stands for the current head d65e7479ad284763d4d640a4f13458e7db67eea8.
· branch feat/usage-tracking-claude-json

…se a bare key on duplicate rows

Three defects the automated reviewers caught, two of them reintroducing at a
different scope the exact failure this PR exists to fix.

score_basis was decided over the whole denominator, before competing was
filtered from it. A bundled, name-only, or disable-model-invocation skill can
carry real native usage while being excluded from the contest entirely, so its
score flipped the basis to native-counters while every actual contender sat at
zero. Those rows then got confidence "inferential", which is a catalog ordering
presenting itself as usage-informed. That is this report's own headline defect,
one scope up. The basis is now decided from the competing rows only.

The bare-key resolver collapsed owners into a set, so two marketplaces shipping
the same plugin, which the report already marks ambiguous-attribution, produced
one distinct qualified name. A bare key then passed the single-owner test and
was reported on BOTH rows, inventing usage for an attribution the audit knows it
cannot make. Owners are counted per entry now.

The markdown renderer still said descriptions are dropped "least-invoked-first"
and described the order as a likelihood band regardless of basis. It now says
lowest-score-first, names decay-weighting and description length as inputs, and
on an unscored run states plainly that the order carries no starvation
information rather than hedging it as inferential.

The 0.39.0 changelog entry also omitted the first-fit, floor and tiebreak
changes entirely.

Suite 96/96 with two regression tests, ruff clean, markdownlint clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
@claude

claude Bot commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 58s —— View job


Security review — /review:security-review

  • Gather context (PR diff since last reviewed head, changed files)
  • Invoke /review:security-review skill and follow its procedure
  • Post findings

Result: no security issues found.

The last successful security review on this PR covered head d65e7479ad284763d4d640a4f13458e7db67eea8. One commit landed since then, bringing HEAD to c96caec (fix(claude-ops): decide the score basis from the contenders, and refuse a bare key on duplicate rows):

git diff --stat d65e747..c96caec
 plugins/claude-ops/CHANGELOG.md                                          | 18 +++++-
 .../skills/audit-skill-visibility/scripts/audit_skill_visibility.py      | 64 ++++++++++++++++------
 .../skills/audit-skill-visibility/scripts/test_audit_skill_visibility.py | 62 +++++++++++++++++++++
 3 files changed, 126 insertions(+), 18 deletions(-)

This is a real code change (not skip-gate-eligible on docs alone — two of the three previously-open code-review findings on this PR are fixed here), so I read the full delta directly:

  • resolve_event_keys — leaf_owners changes from dict[str, set[str]] to dict[str, list[str]] (line 175), so duplicate qualified names (two marketplaces shipping the same plugin) no longer collapse to a single owner; a bare-leaf event against a duplicated name now correctly falls into ambiguous (sorted(set(owners)), line 194) instead of being silently attributed to both rows. Pure in-memory dict/list bookkeeping over already-collected local data — no new I/O, subprocess, or trust-boundary surface.
  • compute_listing — score_basis is now computed from competing (line 873-875) after that list is built, instead of from the full rows denominator, so a bundled/name-only/disable-model-invocation skill's native usage can no longer flip the basis to native-counters while every actual contender scores zero. Same class of change: arithmetic/classification over data already parsed earlier in the script, no new inputs or code paths.
  • The _render_markdown change (further down the same diff, around line 1317+) branches on listing["score_basis"] to stop presenting an unscored/alphabetical ordering as an inferential likelihood band. Prose/formatting only.
  • test_audit_skill_visibility.py — updated assertions covering the duplicate-owner and competing-only-basis cases. No security surface.
  • CHANGELOG.md — prose documenting both fixes. No code.

No GitHub Actions workflow files are touched by this commit or by the PR as a whole. No new file paths, subprocess/eval calls, network access, secrets exposure, or trust-boundary changes are introduced — this remains a local, read-only diagnostic script reconciling already-parsed usage data (~/.claude.json skillUsage, the plugin's own JSONL store, OTEL) with no new fields read.

The full-diff review posted earlier on this PR (no security issues found: no Actions workflows touched, no new file paths/subprocess/eval/shell interpolation, no secrets exposure, read-only local diagnostic) still stands for the current head c96caec418bfd043f4963b8667f49dd88617f021.
· branch feat/usage-tracking-claude-json

@kyle-sexton
kyle-sexton merged commit 1f40f16 into main Sep 1, 2026
60 checks passed
@kyle-sexton
kyle-sexton deleted the feat/usage-tracking-claude-json branch September 1, 2026 01:45
kyle-sexton added a commit that referenced this pull request Sep 1, 2026
…rop order as fact (#3551)

No linked issue

`audit-native-overlap` asserted in two places that skill-listing name-only degradation goes "least-invoked-first", sourced to the skills page. #3532 established from the shipped binary that this is wrong, so one skill in this plugin corrected the claim while its sibling still asserted it.

The upstream-claims row named "the drop order changes" as its own recheck trigger, and that trigger had fired. It now records both what the docs say and what the binary does, routes the mechanism to `audit-skill-visibility/reference/listing-scorer.md`, and watches the binary rather than only the page.

The Budget exposure guidance carried the same false mechanism as an OPERATIVE instruction, which is the line the skill acts on when characterizing an over-budget fleet. Correcting only the claims row would have left that live and made the two sections contradict each other. It now says lowest-score-first and states that raw invocation count is not the exposure ranking and that description length matters.

The 1,536-character per-entry cap and the 1% listing budget in that row are unchanged and still hold; the settings page is authoritative for those and matches.

Found by a post-merge verifier sweeping merged main for the stale wording after #3532.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

audit-skill-visibility: starvation band is alphabetical, and bare usage keys are dropped

1 participant