Repository navigation
The Rust kernel was September's, and nothing said so - #255
Merged
Merged
Conversation
backend-native's published JAR carries classes only -- zero META-INF/models/native entries -- so the .so comes exclusively from a separate platform JAR. The worker fetched that under the fixed name models-kernels-linux-x86_64.jar, and the object on S3 under that name was built 2026-09-28 while the Java side was 0.3.56. Every run in this campaign therefore measured a September Rust kernel against a current library. Both declare abi=6, so it loaded cleanly and there was no symptom at all: released 0.3.56 is sha256 064cfae7... at 485960 bytes, the one in use ce25a986... at 480536. The reports recorded native-kernel-abi=6 and looked entirely consistent, because ABI compatibility is not identity. This is the same stale-hardcode defect I had just fixed for the payload and BACKEND_VERSION in the previous commit, in the same file, two lines away. I fixed two of the three and shipped the third, and it was the one whose staleness nothing downstream could detect. QUAL_KERNELS is now required alongside QUAL_PAYLOAD and QUAL_BACKEND_VERSION, and is checked against the version label by the same rule, so a versionless models-kernels-linux-x86_64.jar is refused rather than silently preferred. The worker also logs the .so's sha256 out of the JAR's own native.properties, because that digest is the only thing distinguishing two ABI-6 kernels and a report should be checkable against the kernel that produced it. Where to get the right one: the release run publishes the platform JARs as artifacts named release-native-<platform>, so the released kernel is downloaded from the release rather than rebuilt and hoped to match. Guard test extended: the versioned kernels jar is accepted, the versionless and the 0.3.54 ones are rejected, all three required inputs are asserted present in the worker, and the classpath is asserted to use $KERNELS rather than a fixed name -- because a fixed name is how this happened.
I fixed the kernels JAR in qual-worker-two-arm.sh and did not look for other copies. parity-worker.sh had it as well, on two lines, identically: an s3 fetch and a classpath entry both naming models-kernels-linux-x86_64.jar literally. Fixing one of two is how the original defect survived in the first place. parity-worker.sh now requires PARITY_KERNELS and logs the .so digest, matching the two-arm worker. The S3 object under the fixed name pointed at the 2026-09-28 build. It now serves the released 0.3.56 kernel, sha256 064cfae7..., so nothing that still reads that name can pick up a stale artifact. The September build is kept under archive-precampaign/ named by its commit, because it is the record of what was actually deployed and deleting it would erase that. And a check so this class cannot be reintroduced by hand: no script in scripts/fleet may name a payload or kernels artifact literally in an s3 fetch or on a classpath. The first version of that check was useless and I nearly shipped it. Its character class was [a-z0-9_]+, which cannot match models-kernels-linux-x86_64.jar because the name contains hyphens, so it passed while the defect was present -- worse than no check at all. It is now verified to fail on reintroduction in BOTH workers and to pass on a clean tree, and that verification ran before the commit rather than after.
… is requalified Measured, one box, paired: same payload, same model, two kernels, each arm run twice in opposite order so first-run JIT and page-cache cost cancels. The kill criterion was satisfied -- the arms are proven distinct by the .so digest each JAR declares, ce25a986 against 064cfae7. Numerics are identical. totalOutputTokens, correctAnswerRate and modelAnswerRate match across all four runs of each model and performanceTier is PRODUCTION_READY in all eight. The campaign's correctness findings stand, including the published nulls in the max-output-tokens sweep, which turn on modelAnswerRate and truncatedAnswerRate and not on speed. On the two inputs the relative gate actually reads -- decodeThroughputRatio against a 0.80 floor and endToEndLatencyRatio against a 1.50 ceiling, prefill being no part of it -- the worst case against us is decode -0.05% and end-to-end p95 +1.76%. Every QUALIFIED verdict has more headroom than that. The tightest on both axes is eurollm Q4_K_S at 1.4% above the decode floor and 3.2% below the end-to-end ceiling. No verdict changes sign. So the twelve vetted models need re-running for provenance, because a report has to name a released version to land, but not to correct a number. Two things recorded rather than smoothed over. The experiment never exercised the Q5_1 half of the kernel delta: both test models carry no Q5_1 tensors, because Q5_K_M is a scheme name and does not imply Q5_1 blocks, and I picked one of them specifically to isolate Q5_1. That error does not change the conclusion only because none of the thirteen pending models carries Q5_1 either, which was verified by range-fetching all thirteen headers rather than assumed from their names. The Q5_1 path stays untested and any future artifact carrying it is outside what this result covers. And the reason none of this was visible: both kernels declare abi=6, so the stale one loaded cleanly and every report recorded a consistent abi. ABI compatibility is not identity, and nothing recorded the .so digest. Both workers now require the kernels JAR by name, log that digest, and are covered by a check that fails if any fleet script names a measurement input literally.
Two checks, each named for a defect that shipped, plus one that proves the others work. check-evidence-references.py: every reference to benchmark-results/ or experiments/ in our own documents must resolve. The 0.3.56 CHANGELOG cited benchmark-results/2026-10-08-q4-1-exact-path/NOTES.md, a directory that never existed, and only a manual look caught it. Scope is deliberately narrow -- those are the paths that back measurements, and a check that reports things nobody can fix gets ignored, which is worse than none. It found seven pre-existing broken references, including a REPRODUCE doc promising five artifacts it does not contain. Those are enumerated in a baseline with reasons; the baseline may shrink and cannot grow, and an entry that starts resolving also fails so it cannot hide a later break. mutation-check.py is the one that matters. It runs each guard twice: clean, where it must pass, and with its own defect injected into a throwaway copy, where it must fail. The reason it exists is that a guard I wrote to catch fixed-name kernel references used [a-z0-9_]+, which cannot match models-kernels-linux-x86_64.jar because the name has hyphens -- it passed while the defect was present, and I nearly committed it. It paid for itself immediately by finding three more defects in guards that looked green: derive.py trusted its committed backend-delta record whenever git could not answer, so on a shallow clone a tampered record would admit a pair whose arms differed on the measured path; the tamper detection only worked when git did. Now fail-closed, with an explicit env var to accept the record on trust. With that fixed, derive.py exited 0 having rejected every pair. A derivation that produced no numbers was reporting success. Now non-zero on zero admissible pairs and on any rejection. And its NOTES check passed vacuously against zero rows, so it would have blessed any prose at all. Now non-zero when there is nothing to check against. All three are one shape: a check reporting green because it had nothing to check. None was visible from a clean run. Adding a guard now means adding its mutation. If you cannot write a mutation that makes your guard fail, the guard does not work yet.
… it working scripts/integrity/README.md quoted the fabricated changelog path verbatim as an example of the defect, and the check cannot tell a citation from an illustration, so it flagged it. CI caught that, not me. Fixed by not writing a literal non-existent path into a document at all: the directory is named without the benchmark-results/ prefix, which reads the same and matches nothing. The alternative was a marker comment the author has to remember, and a guard whose correctness depends on being remembered is the kind that quietly stops working. The rule is now written down in that README under its own heading, because the next person to document a bad path will hit exactly this.
…not a forgery I had called the changelog defect "fabricated", in a mutation name and in a pull-request body. That word is wrong and it is damaging in a permanent public record for a commercial catalog. Wrong, because it describes manufacturing false evidence to deceive. What happened is that a changelog pointed at benchmark-results/2026-10-08-q4-1-exact-path when the real directory is dated 2026-10-09-q4-1-support. The measurements were real, taken on the harness, and every number in that entry traces to the actual NOTES.md -- re-verified in this pass, 11 of 11. A wrong reference to real evidence is a careless citation, not a forgery. Damaging, because this repository exists partly to keep that distinction sharp. Using the word for an incorrect path flattens it against the failure the house rules were written for, and anyone reading the history later cannot tell which one occurred. The mutation is now evidence-references/unresolvable-path, which is what it injects and what the guard detects. The pull-request body that used it has been corrected to "unbacked", which is the gate's own verdict name. One pushed commit message still carries the word. It is not being rewritten: the branch is shared, and a correction that is visible in the history is better than one that edits the record to look as though nothing happened.
derive.py verifies that a paired measurement's two arms came from builds whose delta misses the measured path, which needs the commit range. On CI's shallow clone it cannot, so once it was made fail-closed it went red on a clean tree and the mutation check correctly refused to draw any conclusion from it. fetch-depth: 0 rather than letting it trust a record it cannot verify. That trust is the hole that ran a September kernel for a whole campaign with no symptom.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The defect
backend-native's published JAR carries classes only — zeroMETA-INF/models/nativeentries — so the.socomes exclusively from a separate platform JAR. The worker fetched that under the fixed namemodels-kernels-linux-x86_64.jar, and the object on S3 under that name was built 2026-09-28 while the Java side was 0.3.56.So every run in this campaign measured a September Rust kernel against a current library.
There was no symptom. Both declare
abi=6, so it loaded cleanly and the reports recordednative-kernel-abi=6and looked entirely consistent. ABI compatibility is not identity:064cfae7…ce25a986…How it got shipped
This is the same stale-hardcode defect I had just fixed for
QUAL_PAYLOADandQUAL_BACKEND_VERSION— in the same file, two lines away. I fixed two of the three and shipped the third, and it was the one whose staleness nothing downstream could detect.The fix
QUAL_KERNELSis now required alongside the other two and checked against the version label by the same rule, so a versionlessmodels-kernels-linux-x86_64.jaris refused rather than silently preferred.The worker also logs the
.so's sha256 out of the JAR's ownnative.properties, because that digest is the only thing distinguishing two ABI-6 kernels — a report should be checkable against the kernel that produced it rather than trusted.Where to get the right one: the release run publishes platform JARs as artifacts named
release-native-<platform>, so the released kernel is downloaded from the release rather than rebuilt and hoped to match.Guard test
That last assertion exists because a fixed name is how this happened.
Consequence for the pending work
The 12 vetted models and the 6-model slate have not been run yet, so no landed evidence is affected. The confirmation run will be the first to use the released kernel, which is the point of waiting for the release at all — previously it would have measured released Java against a September
.soand called it 0.3.56.