Work in progress. The instrument is built and tested. No results have been produced with it yet. Nothing here is a finding.
You write a CLAUDE.md — a file of instructions that gets prepended to every request you make
to a coding agent. "Be concise." "Don't hedge." "Answer first, explain after."
Does it work?
Nobody actually knows. You can read the model's replies and form an impression, but an impression can't tell you whether the file changed anything or whether you just got used to it. And the moment you compare two versions of the file, you're comparing two piles of prose by feel.
cc-bench measures it. Swap the config, run the same prompts, and count what changed in the output.
Not "is the answer good". Counters, over the text and over the tool calls:
| verbosity | words per answer, sentence length |
| formatting | bold, headers, lists, tables |
| hedging | "might", "possibly" — and separately, "that said", "on the other hand" |
| answer offset | how far into the reply the first real claim appears |
| shape | does it ask before acting, restate the question, offer a menu, close with a question |
| agentic | tool calls per turn, reads before the first edit, files touched vs files asked for, does it verify after editing |
Every counter is objective, free to compute, and reproducible. There's no model in the loop grading anything, and no ceiling to saturate.
The first version of this project scored configs on a rubric. It was junk, for a reason worth stating publicly: the rubric was written in the vocabulary of the config it was testing. The axes were named after the things the config asked for, so the config won by construction, and there was no axis it could lose. A benchmark that cannot produce a negative result is not measuring anything.
Counting sidesteps that. A counter doesn't know what you were hoping for.
The trade is that counters are blind to correctness — a config that makes a model terse and wrong fingerprints beautifully. So exactly one judged axis survives, task success, and it is reported beside the fingerprint rather than as the headline.
Does my config do anything? Run it against an empty config. If nothing moves, your file is decoration.
Which of my two drafts is stronger? Run them head to head on the same prompts.
Can a cheaper model do what my expensive one does? Measure the behavior you like, save it as a target profile, and see which model-plus-config combination comes closest. This is the part nobody else is doing — existing work measures how far a model drifts from its own trained register, not how close it lands to a target you wrote.
Zero dependencies. Node 22+, ESM, no install step.
git clone https://github.com/thatmike1/cc-bench && cd cc-bench
cp variants.example.json variants.json # declare which configs to compare
node scripts/build-configs.mjs # materialize them
node scripts/run-bench.mjs --engine claude --variants ungoverned,terse
node scripts/report.mjs results/<runId>The example manifest is runnable as shipped — it compares two example configs against an empty
one. To test your own file, point source at it and declare which sections to ablate:
{
"source": "~/.claude/CLAUDE.md",
"variants": {
"current": { "drop": [] },
"no-style": { "drop": ["## Response Style"] }
}
}That's the ablation the tool is built for: your real config against your real config minus one section, so the delta is attributable to that section and nothing else.
variants.json is gitignored. Your personal config never enters the repo.
claude (claude -p), codex (codex exec), and claudex (GPT running inside the Claude
Code harness, so the scaffold is held constant while the model changes).
Your live
~/.claude/CLAUDE.mdis never written.CLAUDE_CONFIG_DIRand aHOMEoverride are both dead ends — the real file loads regardless, and they redirect your credentials on top of it — so eachclaudespawn runs inside a bwrap mount namespace with the variant bound over that one path. The mount dies with the process, interactive Claude Code keeps seeing your own config, and two batches can run different variants at once. bwrap is Linux-only; where it is missing the runner falls back to--isolation swap, which does overwrite your live file for the batch (timestamped backup, restore on exit) — under that mode only, don't use Claude Code interactively while a batch runs. Either way the rest of~/.claude(projects/, history) stays shared, as always.
Three sections, in increasing order of how much weight they carry:
- Headline table — pooled means, no error bars. Orientation only.
- Per-cell variance — spread within each (config × prompt) cell.
- Paired comparison — per-prompt deltas, bootstrap CIs, permutation p-values. This is the only table conclusions come from.
Four things this benchmark does to keep itself honest, each of which took a wrong version first:
Paired on prompts. Prompts differ in verbosity far more than configs do. Comparing raw means throws that variance into the error term and drowns the effect you're looking for.
A pre-registered primary counter set. A wide counter set is a multiple-comparison tax: every counter you add inflates the corrected p-value of every other one, so improving the instrument destroys the evidence. Four counters are registered in advance and carry claims. The rest are computed, reported, and explicitly labelled exploratory — hypothesis-generating, never evidence.
A ~20% relative floor. Smaller effects are not detectable at any run size worth paying for. Design configs for large behavioral changes, or don't report the small ones.
A Goodhart guard. Counters are trivially gameable — "never write the word caveat" zeroes a counter with no behavioral change at all. That's arguably fine, since gaming the counter is obedience, which is what's being measured. It's only a problem when the named counters are the only ones that move. So the tool derives, per run, which counters a config could have targeted directly and which it could only move indirectly, and compares movement across the two. The rule is public; the per-counter membership is not printed by default, because a human iterating against the report will optimize whatever the report shows.
"Nobody benchmarks persistent-config obedience" was this project's original pitch. It's false, and the honest version of the claim is narrower.
| work | what it establishes |
|---|---|
| McMillan, arXiv 2605.10039 | 1,650 Claude Code sessions varying CLAUDE.md size, position, architecture, contradictions. Null on all four. The real effect is within-session decay: ~5.6% lower compliance odds per generated function |
| ContextEcho, arXiv 2605.24279 | Persona drift measured by a judge-free fingerprint. 17 of 23 models exceed |Δ|≥0.30. Compaction does not reset drift; a single-shot anchor can restore it |
| OctoBench, arXiv 2601.10343 | Scaffold-aware instruction following over trajectories, 7,098 objective checklist items |
| Gloaguen et al., arXiv 2602.11988 | Context files do not generally improve task success, and add >20% inference cost |
What's left unoccupied, and where this sits: swapping competing style configs and measuring the behavioral delta between them, and measuring distance from a user-authored target profile.
Counters are also borrowed rather than invented wherever a published operationalization exists — verbosity and formatting follow LMArena's style control so the numbers sit on the same scale as a public leaderboard, and the hedging and praise lexicons come from released, human-validated sets.
Built: the runner, the fingerprint, the fixture, the statistics, the multi-turn drift arm, the Goodhart guard, console and HTML reports, a web control panel.
Not built: the target-profile mechanism, which is the part that makes this a config-authoring tool rather than a config-comparing one.
Not done: a single real run under the current design. The one full run that exists used the retired prompt set and the old counters, and it was inconclusive — the largest raw effect didn't survive correction. That run is kept as a sizing guide and nothing more.
Moving a counter toward a target has not been shown to make a model better. Stylometric features separate human from machine text at near-perfect accuracy while saying nothing whatsoever about whether the text is correct or useful. This tool measures whether your config did what it said. Whether what it said was a good idea is still your problem.