Skip to content

No PostToolUse hook has completed on this machine since 2026-08-12 (formatters silently not running) #3549

Description

@kyle-sexton

Symptom

No PostToolUse hook on this machine has produced a completion or an observable effect since 2026-08-12T09:07:17Z. Nine distinct hooks across six plugins stopped within a roughly ten-hour window that day and have produced nothing in the nineteen days since:

skill-usage-audit, markdown-format, typos-format, eol-normalizer, biome-format, bash-format, cli-flag-verify, stale-path-verify, skill-reference-verify, actionlint-check.

Adjacent events are healthy over the same period. PreToolUse logged 2,404 telemetry records today; PostToolUseFailure logged 7. Every PostToolUse dispatch today ends in a hook_cancelled transcript attachment, with zero hook_success records for any PostToolUse:* hook.

The practical consequence is not limited to telemetry: the format-on-write hooks are silently not running, so files land unformatted with no user-visible error.

Reproduction

Minimal and self-contained. No plugins needed, roughly two minutes. Add to .claude/settings.json, restart, then Write any file:

{"hooks":{
  "PreToolUse":[{"matcher":"Write","hooks":[{"type":"command","command":"date -u +pre-%FT%TZ >> /tmp/hookprobe.log"}]}],
  "PostToolUse":[{"matcher":"Write","hooks":[{"type":"command","command":"date -u +post-%FT%TZ >> /tmp/hookprobe.log"}]}]}}

Expected: one pre- line and one post- line. Observed here: pre- only. Corroborate in the transcript (~/.claude/projects/*/*.jsonl): the PostToolUse:Write record is {"type":"hook_cancelled",...} with no hook_success counterpart.

Two independent effect tests reached the same place:

  • A live Skill tool call produced no store row, no telemetry envelope, and nine hook_cancelled / PostToolUse:Skill records.
  • A .md file written with trailing whitespace and a misspelling was still unformatted two minutes later, with a hook_cancelled / PostToolUse:Write record for that exact toolUseID.

What is ruled out

  • Kill switch. No skill_usage* key in any of the four settings scopes or ~/.claude.json; the manifest default is true.
  • Registration drift. All four installed plugin-cache versions register PostToolUse matcher Skill.
  • Hook correctness. Fed a synthetic payload directly, the hook writes correct rows in both a normal checkout and a linked git worktree.
  • Timeout as the whole story. markdown-format measures 8.2s inside a 15s budget and is equally dead.

Boundary, stated conservatively

  • Last known good: 2026-08-12T09:07:17Z. First confirmed bad: 2026-08-31. The window cannot be narrowed, because no telemetry exists inside it and that absence is the symptom.
  • No version can be pinned. Transcript retention only reaches 2026-08-23. Versions observed since, all bad: 2.1.238, 2.1.241, 2.1.243, 2.1.245, 2.1.251, 2.1.252. This is "not reproducible against a known-good build on this machine", not "regressed in version X".
  • One machine, unconfirmed elsewhere. No second host was tested. This is a report, not an established product defect.
  • No exact upstream match. The nearest, Plugin hooks.json: command hooks silently dropped for PreToolUse/PostToolUse events anthropics/claude-code#34573, is closed and predates this; open candidates (#63047, #42336) are Desktop/Cowork-scoped rather than CLI.

Recorded dissent

A fresh-context verifier independently reproduced the fleet-wide finding and then disputed the disposition, arguing for a two-layer verdict: upstream dispatch failure plus a repo-owned fix, raising skill-usage-audit's declared 5s timeout in hooks.json, on the grounds that it measured 15.4 to 37.6s over 6 of 6 runs while its sibling tool-failure-audit at the same budget came in at 0.3 to 2.9s. Its argument: "If Claude Code's PostToolUse dispatch were fixed tomorrow, this hook would still fail on this machine."

That fix is not being made, for two reasons. It is unfalsifiable at any value while no PostToolUse hook completes at all, and the measurements are load-contaminated: they were taken while bash -c 'exit 0' itself cost 1.1s on this host, and the same hook's two recorded historical runs on the same machine and code were 1181ms and 2261ms, comfortably inside 5s. The durable part of that observation is filed separately as a spawn-count performance finding.

The verifier also corrected three claims in the original investigation, all accepted:

  1. An argument that hook_cancelled records lacking timedOut proved abort rather than timeout was invalid. The binary's wrapper rebuilds the attachment unconditionally and discards that field, so it is absent for every cause. The conclusion survives on the telemetry cutoff; that particular reasoning does not.
  2. "Cancels them before they run" overstates the evidence. tool-failure-audit wrote 7 rows today at timestamps with no corresponding cancellation, so a hook can execute despite a recorded cancellation. The defensible claim is "produces no completions".
  3. A "worktree is clean" claim was false when stated (eight paths, from concurrent sibling agents rather than the hook work).

A ConfigChange lead was considered and dropped: five lifetime firings in a 21,609-line log, last at 2026-08-11T23:33:06Z, which is inside the same window rather than before it, and a hook that only fires on config change tells you nothing by going quiet.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-humanHuman-in-the-loop required; autonomous sessions must not resolve items carrying this.priority: highSignificant impact, or blocks an imminent release; staff this cycle.work-class: read-onlyAudits, research, reports. No repository mutation; tracker and queue writes only.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions