Skip to content

Write up the multi-signal language-detection architecture as Claim 11 #3119

Description

@squid-protocol

The detection cascade in gitgalaxy/standards/language_lens.py is one of the more architecturally distinctive parts of the engine, and it has no argued writeup anywhere.

What exists: docs/wiki/02-05-language-lens.md — 52 lines, in the 02- pipeline-subsystem series. It describes the tiers accurately but makes no case and draws no comparison. That is the right job for a mechanical reference page and it should keep doing it.

Where argued positions live: the 03- series, Claims 1-10. docs/wiki/03-10-claim-10-ast-vs-heuristic-parsing.md does exactly this job for the AST-vs-heuristic decision (intersection / our advantage / their advantage / honest limits). Claim 11 is free — the series currently runs 03-01 through 03-10, with docs/wiki/03-20-future-outlooks.md after it.

Deliverable: docs/wiki/03-11-claim-11-<slug>.md, added to the mkdocs.yml nav and docs/wiki/index.md (both already index the claims series). docs/document_alignment_guide.md also references claims and may need a line.

docs/wiki/02-05-language-lens.md stays as the mechanical reference and should gain a cross-link to Claim 11, and vice versa — not replaced, not merged.

Content outline

1. Thesis

File extensions are an assertion, not evidence. Detection should be a multi-signal inference with an explicit trust level attached to the answer, not a waterfall that returns a bare label.

2. The evidence sources actually combined

All zero-dependency — no trained model, no sample corpus:

  • exact filename anchors;
  • extension, with multi-suffix unwrapping for wrapper extensions (script.sh.template -> .sh);
  • shebang;
  • prose anchors (README/LICENSE-family, with affix matching);
  • sibling anchors — foo.h resolves to C when foo.c is in the tally, Obj-C when foo.m is;
  • ecosystem gravity — directory-local extension tally plus per-language discriminators and disqualifiers, gated on ECOSYSTEM_DOMINANCE_MIN (0.70);
  • internal_discriminator content regexes, consulted only for registry-declared collision extensions;
  • a full lexical scan scoring the file against each candidate language's own structural-signature rules;
  • a discovery funnel for extensionless files that isolates comment family first, then requires a 1.5x margin over the runner-up;
  • caller-supplied context priors.

3. The three genuinely differentiated properties

This is the honest core of the claim. The doc should say exactly this much and no more.

  • Agreement-based tiering with provenance. Two independent signals concurring yields Tier 0 at 0.999; a single indicator yields Tier 2 at 0.91. Every result carries a lock_tier and a human-readable source_proof — the engine reports not just the answer but how it was reached and how much to trust it. Mainstream detectors return a label.
  • Identity conflict as a refusal, and as a security signal. When extension and shebang disagree, the file is not resolved to a best guess: it drops to Tier 5 ("Absolute Distrust"), undeterminable, intensity 0.0, with an Identity Masking anomaly flag. The contradiction is the finding. This is the piece with no counterpart I am aware of in Linguist, enry, guesslang, Pygments or libmagic.
  • Repo-context inference. Gravity and sibling anchors classify a file using its neighbours, which stateless file-by-file classifiers structurally cannot express. This is what makes the mainframe family tractable: an EQU-only copybook is HLASM because it sits beside JCL and COBOL.

4. Where it is convergent rather than novel

State this plainly — it makes the doc credible rather than weaker.

The collision registry + content-discriminator mechanism is functionally what GitHub Linguist does in heuristics.yml. The per-language lexical self-scoring is architecturally what Pygments does with analyse_text().

Arriving independently at the same disambiguation structure as the tool calibrated against millions of repositories is a validation of the design, not an embarrassment. But a writeup that claims novelty for those two parts is falsifiable in about five minutes by anyone who knows the field, and would discredit the three real claims above by association.

The honest framing: independently arrived at the field's disambiguation architecture, with no dependencies and no training data, then added provenance, refusal, and repo-context on top.

5. Honest limits

Mirroring Claim 10's own §3, which is why that doc reads as credible. The lexical scan's scoring constants are hand-tuned and uncalibrated — verified in _tier_3_lexical_scan:

  • per-hit weights, and a comment-delimiter bonus of += 15.0 awarded on a plain substring test (if d in content);
  • two 1.25x boosts (language_lens.py:772, :778);
  • a hardcoded 0.4x handicap for exactly three broad-ruleset languages — if lid in ("abap", "fortran", "cobol") (:789);
  • log-normalisation by LOC (raw_score / math.log1p(loc));
  • a confidence derived by dividing by 50 (confidence = min(top_signal / 50.0, 1.0), :802).

Plus: first-match-wins over registry order arbitrates ties (#3118), and the two defect classes found in the same review — the EQU-copybook gravity miss (#3110) and the unanchored-shebang Tier 5 false conflicts (#3116).

A claims doc that omits these while they are open issues in the same repo would not survive a skeptical reader.

6. Blocked on a number

The doc's weakest possible form asserts architectural superiority with no measured accuracy. Its strongest form opens with a measured per-language precision/recall figure on a real labeled corpus.

#3117 (detection-accuracy harness) is a soft dependency: the writeup can be drafted now, but the headline claim should not be published until there is a number to put in it.

Related

#3117 (the measurement this is blocked on), #3116 and #3110 (the defects the honest-limits section must disclose), #3118 (the registry-order caveat).

Metadata

Metadata

Assignees

No one assigned

    Labels

    core-engineModifications to the central physics and parsing enginedocumentationUpdates to standard operating procedures, guides, or wikipriority: lowUI tweaks, documentation, and minor optimizations

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions