Skip to content

docs(wiki): Claim 11 — multi-signal language detection (#3119) - #3193

Merged
squid-protocol merged 1 commit into
mainfrom
docs/3119-claim-11
Sep 19, 2026
Merged

squid-protocol merged 1 commit into
mainfrom
docs/3119-claim-11

Conversation

@squid-protocol

Copy link
Copy Markdown
Owner

Closes #3119.

What

Writes up the multi-signal detection cascade in language_lens.py as Claim 11 in the 03- series — the argued-position job that 02-05-language-lens.md (mechanical reference) deliberately doesn't do. New page: docs/wiki/03-11-claim-11-multi-signal-detection.md.

Follows Claim 10's intellectual-honesty structure:

  1. Thesis — a file extension is an assertion, not evidence; detection is inference with an explicit trust level, not a waterfall returning a bare label.
  2. The signals combined — all zero-dependency (exact-filename, extension w/ wrapper unwrapping, shebang, prose anchors, sibling anchors, ecosystem gravity, internal_discriminator, lexical scan, discovery funnel, context priors).
  3. The three genuinely differentiated properties — agreement-based tiering with lock_tier + source_proof provenance; identity-conflict-as-refusal (Tier 5 + Identity Masking security flag); repo-context inference (incl. the Ecosystem gravity should weigh sibling classifications, not extension counts #3137 sibling-classification vote — an EQU-only copybook is HLASM because of where it sits).
  4. Where it is convergent, not novel — states plainly that the collision-registry mechanism is functionally Linguist's heuristics.yml and the lexical self-scoring is Pygments' analyse_text(); frames independent arrival as validation, not invention.
  5. Honest limits — the hand-tuned, uncalibrated lexical constants (+= 15.0 on a substring test, two 1.25× boosts, the 0.4× abap/fortran/cobol handicap, log1p normalisation, /50 confidence — all verified in source), plus the four harness-found defects (Detection hardening: EQU-only .asm copybooks misclassify as assembly (hlasm discriminator gap) #3110, Shebang matching is unanchored substring search: canonical tclsh/wish/ts-node shebangs trigger false Identity Conflicts (Tier 5) #3116, Pin internal_discriminator and shebang resolution order for contested extensions #3118, Ecosystem gravity cannot resolve a same-extension collision: the contested extension votes for one of its own claimants #3132) the engine has since fixed.
  6. The number — §6 of Write up the multi-signal language-detection architecture as Claim 11 #3119 soft-blocked this on a measured figure (Add a language-detection accuracy harness (per-language precision/recall + confusion matrix + committed baseline) #3117). That harness now exists and, after Ecosystem gravity cannot resolve a same-extension collision: the contested extension votes for one of its own claimants #3132/Ecosystem gravity should weigh sibling classifications, not extension counts #3137, reports 1.0 determinable (3,231/3,231), 509/509 on the hand-labelled contested subset, and an empty confusion matrix — with the extension-circularity and "generalisation estimate, not universal proof" caveats stated inline.

Wiring

  • mkdocs.yml nav (between Claim 10 and Future Outlooks)
  • docs/wiki/index.md — both the Systems-Engineers list and the numbered claims list
  • docs/document_alignment_guide.md — new claim-map row + the "10 → 11 numbered Claims" count
  • Reciprocal cross-link with 02-05-language-lens.md (and fixed that page's broken back-to-index footer)

Validation

mkdocs build --strict shows zero warnings on any of the touched/added files; the new page and all its links resolve. (The build's 8 pre-existing warnings are unrelated dead ../vectors.md / ../zero_dependency_mode.md links in other pages, present before this PR.)

Docs only — no engine change.

🤖 Generated with Claude Code

Writes up the detection cascade in language_lens.py as an argued Claim page
(the 03- series), the job 02-05-language-lens.md (mechanical reference)
deliberately does not do. Follows Claim 10's intellectual-honesty structure:
thesis (extension is an assertion, not evidence), the zero-dependency signals
combined, the three genuinely differentiated properties (provenance tiering,
identity-conflict refusal as a security signal, repo-context inference incl.
the #3137 sibling-classification vote), an explicit 'where it is convergent
not novel' section (Linguist heuristics.yml, Pygments analyse_text), and
honest limits (the hand-tuned uncalibrated lexical constants, and the four
harness-found defects #3110/#3116/#3118/#3132 now fixed).

#3119's headline was soft-blocked on a measured number (#3117); that harness
now exists and, post #3132/#3137, reports 1.0 determinable (3231/3231),
509/509 on the hand-labelled contested subset, and an empty confusion matrix
-- with the extension-circularity and generalisation-estimate caveats stated.

Wired into mkdocs.yml nav, docs/wiki/index.md (both listings), the
document_alignment_guide claim map (and its '10 -> 11 claims' count), and a
reciprocal cross-link with 02-05 (also fixing that page's broken index footer).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

🐦‍⬛ Muninn Security Scan

✅ No security issues found.

🐦‍⬛ Powered by Muninn · Skald Lab

@squid-protocol
squid-protocol merged commit dacfec5 into main Sep 19, 2026
28 checks passed
@squid-protocol
squid-protocol deleted the docs/3119-claim-11 branch September 19, 2026 01:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Write up the multi-signal language-detection architecture as Claim 11

1 participant