Found by a design review of the language-detection cascade (gitgalaxy/standards/language_lens.py).
The gap
This repo measures extraction accuracy and cross-language measurement consistency, but nothing for detection correctness.
What exists today:
| surface |
measures |
artifacts |
| extraction accuracy |
per-language function/class/arg fidelity vs tree-sitter |
tests/tools/tree_sitter_accuracy_audit.py, 30 committed tests/tree_sitter_accuracy_baseline_<lang>.json files, the ast-accuracy-audit CI job, docs/self_scan/tree_sitter_accuracy_history.csv |
| measurement consistency |
cross-language rule agreement |
keyword-rosetta's verifier + bias report |
| detection correctness |
nothing |
— |
The golden crucible pins diffs (regression), not accuracy. So a classification can be stably, permanently wrong and every gate stays green.
Two live consequences, both found by manual inspection in one afternoon
Neither is visible to any existing gate. For #3116 I confirmed why: a sweep of the pinned language-crucible corpus for tclsh, wish, jimsh, ts-node, swift-sh, racketsh and micropython shebangs returns zero files for all seven. There is no corpus coverage of the failing shapes, so no golden-master diff can ever fire on them. A harness that measures accuracy against labeled inputs would have surfaced both systematically.
Ask
A tests/tools/detection_accuracy_audit.py in the same shape as the tree-sitter audit: run the real LanguageDetector over a labeled corpus, emit per-language precision/recall and the top confusion pairs, and commit a baseline JSON that CI gates against regressions.
Critical design note — or the harness will produce garbage
The language-crucible's data/<lang>/ directory grouping is NOT per-file ground truth. gitgalaxy/standards/how_to_add_a_language.md says so explicitly (Step 5, item 2):
Filter every diff by the real file extension in its path, not by the crucible corpus's directory-group label — a directory grouped under one language can legitimately bundle files of another (e.g. a Dockerfile-grouped repo containing .go files).
Labeling strategies for the implementer to weigh, deliberately un-adjudicated:
- Extension-as-label for the uncontested extensions, plus hand-labeling only the
COLLISION_FREQUENCIES members. The contested set is 13 extensions on current main (.inc .h .py .cshtml .c .y .m .map .sql .ddl .dml .asm .cmd), so this is tractable and targets exactly the interesting cases.
- A committed per-file label manifest, curated once.
- Restrict scoring to the contested subset and report coverage honestly rather than claiming a whole-corpus number.
Whichever is chosen, the harness must not silently treat the directory name as truth.
Metrics worth emitting beyond precision/recall
All cheap, and uniquely informative for this architecture:
Seed regression cases
Prior art
GitHub Linguist's samples/ tree is the de-facto labeled corpus for this problem, and its CI tests classification against it — worth a look for corpus material. Licensing of any borrowed sample files needs checking before vendoring; do not assume it is compatible.
Why it is worth doing
Beyond catching the two open defects: this converts "our detection architecture is well designed" into a defensible number for prospects, which no current artifact provides.
Related
Siblings from the same review: #3116 (the shebang defect this harness would gate) and #3118 (registry-order pinning — a narrower, cheaper test that covers the collision subset while this harness is built).
Found by a design review of the language-detection cascade (
gitgalaxy/standards/language_lens.py).The gap
This repo measures extraction accuracy and cross-language measurement consistency, but nothing for detection correctness.
What exists today:
tests/tools/tree_sitter_accuracy_audit.py, 30 committedtests/tree_sitter_accuracy_baseline_<lang>.jsonfiles, theast-accuracy-auditCI job,docs/self_scan/tree_sitter_accuracy_history.csvThe golden crucible pins diffs (regression), not accuracy. So a classification can be stably, permanently wrong and every gate stays green.
Two live consequences, both found by manual inspection in one afternoon
ASMCOPY/REGISTRS.asm, an EQU-only HLASM copybook, classifies asassemblyvia Tier 1.5 ecosystem gravity.tclsh/wish/ts-nodeshebangs resolve to Tier 5undeterminablewith a falseIdentity Maskingsecurity flag.Neither is visible to any existing gate. For #3116 I confirmed why: a sweep of the pinned language-crucible corpus for
tclsh,wish,jimsh,ts-node,swift-sh,racketshandmicropythonshebangs returns zero files for all seven. There is no corpus coverage of the failing shapes, so no golden-master diff can ever fire on them. A harness that measures accuracy against labeled inputs would have surfaced both systematically.Ask
A
tests/tools/detection_accuracy_audit.pyin the same shape as the tree-sitter audit: run the realLanguageDetectorover a labeled corpus, emit per-language precision/recall and the top confusion pairs, and commit a baseline JSON that CI gates against regressions.Critical design note — or the harness will produce garbage
The language-crucible's
data/<lang>/directory grouping is NOT per-file ground truth.gitgalaxy/standards/how_to_add_a_language.mdsays so explicitly (Step 5, item 2):Labeling strategies for the implementer to weigh, deliberately un-adjudicated:
COLLISION_FREQUENCIESmembers. The contested set is 13 extensions on current main (.inc .h .py .cshtml .c .y .m .map .sql .ddl .dml .asm .cmd), so this is tractable and targets exactly the interesting cases.Whichever is chosen, the harness must not silently treat the directory name as truth.
Metrics worth emitting beyond precision/recall
All cheap, and uniquely informative for this architecture:
undeterminableplusplaintextfallbacks.Seed regression cases
.asmhlasm-vs-assembly,.cmdrexx-vs-batch,.sqldb2_sql-vs-sqlite,.mmatlab-vs-objective-c,.pypython-vs-embedded_python.Prior art
GitHub Linguist's
samples/tree is the de-facto labeled corpus for this problem, and its CI tests classification against it — worth a look for corpus material. Licensing of any borrowed sample files needs checking before vendoring; do not assume it is compatible.Why it is worth doing
Beyond catching the two open defects: this converts "our detection architecture is well designed" into a defensible number for prospects, which no current artifact provides.
Related
Siblings from the same review: #3116 (the shebang defect this harness would gate) and #3118 (registry-order pinning — a narrower, cheaper test that covers the collision subset while this harness is built).