Skip to content

Ecosystem gravity should weigh sibling classifications, not extension counts #3137

Description

@squid-protocol

Split out of #3132, which is scoped to that bug's symptom (7 misclassified files, 3 now fixed by a content signal in #3136). This is the architectural change the residual 4 need, stated on its own because it has a different risk profile and cannot be done cheaply — every cheap version has been tried and measured.

The structural problem

_evaluate_ecosystem_gravity resolves a neighbourhood by counting extensions. For a contested extension that is information-free by construction, because both rivals are scored over the same files:

  • the base-mass fallback (language_lens.py ~:597) counts the contested extension itself when a candidate has no other-extension neighbour, so python and embedded_python both score on the same 14 .py files;
  • the only thing that then separates them is a discriminators entry — and python lists .py, its own contested extension, which is literally where the reported "72% Local Dominance" comes from (base 14 + 14×2 = 42 against embedded_python's 14 + 1×2 = 16).

Gravity discriminates a mixed-extension neighbourhood well — the .asm-beside-.jcl-and-.cbl case it was built for, and #2504's .cmd routing. It is structurally blind to same-extension rivals, which today are .py (python/embedded_python), .m (matlab/objective-c) and .cmd (batch/rexx).

Three measured disproofs (baseline 0.9984 overall / 0.9920 contested, #3117 harness)

candidate overall contested what breaks
gravity abstains on same-extension collisions 0.9840 0.9034 32 sqlite → db2_sql
exclude the contested ext from every discriminator list 0.9850 0.9095 same 32 sqlite files
remove only python's .py self-reference 0.9964 0.9799 7 plain-python → embedded_python, 2 → plaintext

The self-reference does real work in both directions, so it cannot be removed for either claimant without something to replace it. And for .sql there is nothing available: the strongest sqlite-only content markers cover about a quarter of the corpus's 80 .sql files (AUTOINCREMENT 19, INTEGER PRIMARY KEY 18, PRAGMA 2, sqlite_master/WITHOUT ROWID/dot-commands/USING fts ≈ 0).

The proposal

Score a neighbourhood by what its siblings actually resolved to, not by their filenames. A sibling that classified as embedded_python at Tier 2 via its own content is real evidence about its neighbours; a sibling that merely shares the contested extension is not evidence at all. Sketch, for the implementer to weigh rather than follow:

  1. classify the directory's files that are resolvable without gravity (exact match, uncontested extension, shebang, internal discriminator);
  2. use only those verdicts as the neighbourhood vote, weighting Tier 0/2 content evidence above extension-only locks;
  3. abstain when the resolvable subset is empty or evenly split, letting Tier 3 decide;
  4. keep the existing extension-count path for mixed-extension neighbourhoods, which it handles correctly today.

This makes gravity's own input causally prior to its output, which is the property it currently lacks. Note the ordering problem it must not introduce: gravity runs per file, so "what did my siblings resolve to" needs either a two-pass scan or memoised per-directory resolution — worth deciding explicitly, since the current single-pass design is why extension counting was chosen in the first place.

Acceptance

Not "the 4 files pass." The bar is: no regression on the #3117 harness (overall ≥ 0.9984, contested subset ≥ 0.9920, refusals ≤ 1, Tier 5 conflicts ≤ 1) and the residual embedded_python errors reduced. Re-run tests/tools/detection_accuracy_audit.py --errors on every iteration; three plausible fixes have already regressed, two of them while looking more principled than the one that worked.

Related: #3132 (the symptom and the disproof record), #3117 (the harness this must be measured on), #3136 (the content-signal partial fix), #2504 (the .cmd case gravity resolves correctly today and must keep resolving).

Metadata

Metadata

Assignees

No one assigned

    Labels

    core-engineModifications to the central physics and parsing engineenhancementNew feature, sensor, or structural signature

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions