Repository navigation
Ecosystem gravity cannot resolve a same-extension collision: the contested extension votes for one of its own claimants #3132
Description
Activity
- addedbugUnintended behavior or logic failure in the engineUnintended behavior or logic failure in the enginecore-engineModifications to the central physics and parsing engineModifications to the central physics and parsing enginepriority: highCore feature broken, but workarounds existCore feature broken, but workarounds exist
on Sep 17, 2026 - changed the title
[-]Ecosystem gravity cannot resolve a same-extension collision and votes for the extension_map winner[/-][+]Ecosystem gravity cannot resolve a same-extension collision: the contested extension votes for one of its own claimants[/+]on Sep 17, 2026 Partially fixed in #3136; staying open for the residual 4 files. Three candidate fixes are now measured and all three are net losses — recorded here and as a comment at the gravity scoring line so they are not re-derived.
Fixed (3 of 7 files)
Three of the seven —
ota.py,mpu6050.py,tsl2591.py— do carry a MicroPython-exclusive import thatembedded_python'sinternal_discriminatorsimply did not list. They never needed gravity: a content signal at Tier 2 pre-empts Tier 1.5 entirely, which is the right ordering for a collision. Added themicropythonmodule itself and theu-prefixed reduced stdlib (enumerated, notu\w+, soutils/ui/uuid/unittestcannot trip it).Harness effect: overall 0.9974 → 0.9984, independent contested subset 0.9859 → 0.9920 (490/497 → 493/497). No regressions.
Still open (4 files)
diagnostics.py,logging.py,meowprotocol.py,tester.pyimport nothing MicroPython-specific and are genuinely indistinguishable per-file. That is precisely the case neighbourhood context exists to solve — and precisely where the self-reference bias bites.Three measured disproofs (baseline 0.9984 / 0.9920)
candidate fix overall contested subset what breaks gravity abstains on same-extension collisions 0.9840 0.9034 32 sqlite → db2_sql exclude contested ext from every discriminator list 0.9850 0.9095 same 32 sqlite files remove only python's .pyself-reference0.9964 0.9799 7 plain-python → embedded_python, 2 → plaintext The third is the one the global experiment could not isolate: it leaves sqlite's
.sqlself-reference intact and still loses, because the self-reference does real work in both directions — suppressing embedded_python where python is right, and vice versa.The blocker, now quantified rather than asserted
My earlier note said
.sqlneeds a content signal before the self-reference can go. Measured: the strongest sqlite-only markers cover roughly a quarter of the corpus's 80.sqlfiles —AUTOINCREMENT19,INTEGER PRIMARY KEY18,PRAGMA2, andsqlite_master/WITHOUT ROWID/ dot-commands /USING ftsessentially 0. So aninternal_discriminatorfor sqlite is not a viable replacement on this corpus.What would actually move this
Not a tweak to the scoring. Gravity resolves a neighbourhood by extension counting, which cannot separate two languages that share the extension — the base-mass fallback at
language_lens.py:~597even counts the contested extension for both rivals, making that term identically zero-information. A real fix needs gravity to weigh sibling classifications (what the neighbours actually resolved to, via their own content) rather than their filenames. That is a larger change and it needs the #3117 harness on every iteration; every cheap version of it has now been tried and measured.- added a commit that references this issue
on Sep 17, 2026 - added a commit that references this issue
on Sep 19, 2026
Found by the first run of the detection-accuracy harness (PR #3130, closes #3117) against the pinned language-crucible corpus. Baseline: 3,072 files across 51 languages at 0.9964 overall / 0.9859 on the independently-labelled contested subset. This is the largest single error class in the remaining 11-file error set.
Measured
7 of the 14
embedded_python/meow_turtle/*.pyfiles classify aspython, at Tier 1.5 with proofEcosystem Consensus Lock (72% Local Dominance):diagnostics.py,logging.py,meowprotocol.py,mpu6050.py,ota.py,tester.py,tsl2591.pyReproduced by passing the repo
ext_tallytoinspect()— note gravity is gated onext in self.COLLISION_FREQUENCIES and ext_tally and lock_tier > 2(language_lens.py:452), so a bareinspect()with no tally skips this path entirely and lands on Tier 4 instead.Why Tiers 1 and 2 correctly decline
.pyis a registered collision claimed by bothpythonandembedded_python, so Tier 1 refuses to lock. Tier 2'sinternal_discriminatorforembedded_pythononly fires on a hardware import:Exactly 7 of the 14 files match it (
actuators,app,bldc_driver,boot,pio_programs,sensors,vibration_driver— all classify correctly). The other 7 are the telemetry/protocol/driver-helper layer and import none of those modules. Gravity then decides.Mechanism — measured, and not where you would expect
Gravity does not consult
self.extension_map._evaluate_ecosystem_gravity(language_lens.py:556) gathers every claimant of the extension and scores each by neighbourhood mass. Replaying its arithmetic on this directory (local tally:{'.py': 14}):python.py3 .py2 .pyw .pyi .pyx .pxd ….py→ 1414 + 14*2= 42embedded_python.mpyboot.py→ 114 + 1*2= 1642 / (42 + 16)= 0.7241 — exactly the observed 72%.Two compounding defects produce that:
_evaluate_ecosystem_gravityfalls back to counting the contested extension itself (base_contributors[ext] = tally.get(ext.lower(), 0),:597). Both rivals therefore score on the same 14 files, so the base signal is identically zero-information.pythonlists.pyin its owndiscriminators. So the contested extension is counted a second time, at double weight, as positive evidence for one of its own claimants. That single self-referential entry is what turns a genuine tie into a 72% "consensus".The architectural point stands regardless: gravity can discriminate a different-extension neighbourhood (the
.asm-beside-.jcl-and-.cblcase it was built for) but is structurally blind to same-extension rivals. On this corpus it is not a weak signal, it is a systematically biased one — and the bias direction is set by a discriminator list, not by registry order or map precedence.Affects all three multi-claimant extensions:
.py(python/embedded_python),.m(matlab/objective-c),.cmd(batch/rexx).Measured fix direction
If gravity abstains on a same-extension collision, these files fall through to the Tier 3 lexical scan, which scores content against every candidate's full rule set. Verified directly — running
inspect()without a tally (which bypasses gravity) returns:embedded_pythonfor 5:logging,meowprotocol,mpu6050,tester,tsl2591(Tier 4, proofCollision Resolved (.py -> embedded_python))pythonfor 2:diagnostics,otaSo abstaining takes this class from 0/7 to 5/7 correct, using the tier actually designed to discriminate content. Record the residual 2 as a known limit, not a complete fix.
Worth weighing alongside abstention: removing
.pyfrompython'sdiscriminatorsis a one-line change that addresses defect 2 directly. It is not sufficient on its own — it would leave a 50/50 tie broken bymax()'s iteration order, which is exactly the silent order-dependence #3118 was about — but it is the narrower of the two levers and they are independent.Related
This is #3110's mechanism in a second form: there gravity over-rode a content signal for an EQU-only HLASM copybook, here it over-rides the absence of one. See also #3117 (the harness that found this).