Skip to content

The Rust kernel was September's, and nothing said so - #255

Merged
bsbodden merged 8 commits into
mainfrom
fix/fleet-kernels-jar-required
Oct 10, 2026
Merged

bsbodden merged 8 commits into
mainfrom
fix/fleet-kernels-jar-required

Conversation

@bsbodden

Copy link
Copy Markdown
Member

The defect

backend-native's published JAR carries classes only — zero META-INF/models/native entries — so the .so comes exclusively from a separate platform JAR. The worker fetched that under the fixed name models-kernels-linux-x86_64.jar, and the object on S3 under that name was built 2026-09-28 while the Java side was 0.3.56.

So every run in this campaign measured a September Rust kernel against a current library.

There was no symptom. Both declare abi=6, so it loaded cleanly and the reports recorded native-kernel-abi=6 and looked entirely consistent. ABI compatibility is not identity:

sha256 size
released 0.3.56 064cfae7… 485,960
the one actually in use ce25a986… 480,536

How it got shipped

This is the same stale-hardcode defect I had just fixed for QUAL_PAYLOAD and QUAL_BACKEND_VERSION — in the same file, two lines away. I fixed two of the three and shipped the third, and it was the one whose staleness nothing downstream could detect.

The fix

QUAL_KERNELS is now required alongside the other two and checked against the version label by the same rule, so a versionless models-kernels-linux-x86_64.jar is refused rather than silently preferred.

The worker also logs the .so's sha256 out of the JAR's own native.properties, because that digest is the only thing distinguishing two ABI-6 kernels — a report should be checkable against the kernel that produced it rather than trusted.

Where to get the right one: the release run publishes platform JARs as artifacts named release-native-<platform>, so the released kernel is downloaded from the release rather than rebuilt and hoped to match.

Guard test

  ok   models-kernels-linux-x86_64-0.3.56.jar accepted for models@0.3.56+v25-666bb48c61e0
  ok   models-kernels-linux-x86_64.jar        rejected
  ok   models-kernels-linux-x86_64-0.3.54.jar rejected
  ok   refuses without QUAL_PAYLOAD / QUAL_KERNELS / QUAL_BACKEND_VERSION
  ok   classpath uses the named kernels jar, not a fixed name

That last assertion exists because a fixed name is how this happened.

Consequence for the pending work

The 12 vetted models and the 6-model slate have not been run yet, so no landed evidence is affected. The confirmation run will be the first to use the released kernel, which is the point of waiting for the release at all — previously it would have measured released Java against a September .so and called it 0.3.56.

backend-native's published JAR carries classes only -- zero
META-INF/models/native entries -- so the .so comes exclusively from a separate
platform JAR. The worker fetched that under the fixed name
models-kernels-linux-x86_64.jar, and the object on S3 under that name was built
2026-09-28 while the Java side was 0.3.56.

Every run in this campaign therefore measured a September Rust kernel against a
current library. Both declare abi=6, so it loaded cleanly and there was no symptom
at all: released 0.3.56 is sha256 064cfae7... at 485960 bytes, the one in use
ce25a986... at 480536. The reports recorded native-kernel-abi=6 and looked
entirely consistent, because ABI compatibility is not identity.

This is the same stale-hardcode defect I had just fixed for the payload and
BACKEND_VERSION in the previous commit, in the same file, two lines away. I fixed
two of the three and shipped the third, and it was the one whose staleness nothing
downstream could detect.

QUAL_KERNELS is now required alongside QUAL_PAYLOAD and QUAL_BACKEND_VERSION, and
is checked against the version label by the same rule, so a versionless
models-kernels-linux-x86_64.jar is refused rather than silently preferred. The
worker also logs the .so's sha256 out of the JAR's own native.properties, because
that digest is the only thing distinguishing two ABI-6 kernels and a report should
be checkable against the kernel that produced it.

Where to get the right one: the release run publishes the platform JARs as
artifacts named release-native-<platform>, so the released kernel is downloaded
from the release rather than rebuilt and hoped to match.

Guard test extended: the versioned kernels jar is accepted, the versionless and
the 0.3.54 ones are rejected, all three required inputs are asserted present in
the worker, and the classpath is asserted to use $KERNELS rather than a fixed
name -- because a fixed name is how this happened.
I fixed the kernels JAR in qual-worker-two-arm.sh and did not look for other
copies. parity-worker.sh had it as well, on two lines, identically: an s3 fetch
and a classpath entry both naming models-kernels-linux-x86_64.jar literally.
Fixing one of two is how the original defect survived in the first place.

parity-worker.sh now requires PARITY_KERNELS and logs the .so digest, matching the
two-arm worker.

The S3 object under the fixed name pointed at the 2026-09-28 build. It now serves
the released 0.3.56 kernel, sha256 064cfae7..., so nothing that still reads that
name can pick up a stale artifact. The September build is kept under
archive-precampaign/ named by its commit, because it is the record of what was
actually deployed and deleting it would erase that.

And a check so this class cannot be reintroduced by hand: no script in
scripts/fleet may name a payload or kernels artifact literally in an s3 fetch or
on a classpath.

The first version of that check was useless and I nearly shipped it. Its
character class was [a-z0-9_]+, which cannot match models-kernels-linux-x86_64.jar
because the name contains hyphens, so it passed while the defect was present --
worse than no check at all. It is now verified to fail on reintroduction in BOTH
workers and to pass on a clean tree, and that verification ran before the commit
rather than after.
… is requalified

Measured, one box, paired: same payload, same model, two kernels, each arm run
twice in opposite order so first-run JIT and page-cache cost cancels. The kill
criterion was satisfied -- the arms are proven distinct by the .so digest each JAR
declares, ce25a986 against 064cfae7.

Numerics are identical. totalOutputTokens, correctAnswerRate and modelAnswerRate
match across all four runs of each model and performanceTier is PRODUCTION_READY
in all eight. The campaign's correctness findings stand, including the published
nulls in the max-output-tokens sweep, which turn on modelAnswerRate and
truncatedAnswerRate and not on speed.

On the two inputs the relative gate actually reads -- decodeThroughputRatio
against a 0.80 floor and endToEndLatencyRatio against a 1.50 ceiling, prefill
being no part of it -- the worst case against us is decode -0.05% and end-to-end
p95 +1.76%. Every QUALIFIED verdict has more headroom than that. The tightest on
both axes is eurollm Q4_K_S at 1.4% above the decode floor and 3.2% below the
end-to-end ceiling. No verdict changes sign.

So the twelve vetted models need re-running for provenance, because a report has
to name a released version to land, but not to correct a number.

Two things recorded rather than smoothed over. The experiment never exercised the
Q5_1 half of the kernel delta: both test models carry no Q5_1 tensors, because
Q5_K_M is a scheme name and does not imply Q5_1 blocks, and I picked one of them
specifically to isolate Q5_1. That error does not change the conclusion only
because none of the thirteen pending models carries Q5_1 either, which was
verified by range-fetching all thirteen headers rather than assumed from their
names. The Q5_1 path stays untested and any future artifact carrying it is outside
what this result covers.

And the reason none of this was visible: both kernels declare abi=6, so the stale
one loaded cleanly and every report recorded a consistent abi. ABI compatibility
is not identity, and nothing recorded the .so digest. Both workers now require the
kernels JAR by name, log that digest, and are covered by a check that fails if any
fleet script names a measurement input literally.
Two checks, each named for a defect that shipped, plus one that proves the others
work.

check-evidence-references.py: every reference to benchmark-results/ or
experiments/ in our own documents must resolve. The 0.3.56 CHANGELOG cited
benchmark-results/2026-10-08-q4-1-exact-path/NOTES.md, a directory that never
existed, and only a manual look caught it. Scope is deliberately narrow -- those
are the paths that back measurements, and a check that reports things nobody can
fix gets ignored, which is worse than none. It found seven pre-existing broken
references, including a REPRODUCE doc promising five artifacts it does not
contain. Those are enumerated in a baseline with reasons; the baseline may shrink
and cannot grow, and an entry that starts resolving also fails so it cannot hide a
later break.

mutation-check.py is the one that matters. It runs each guard twice: clean, where
it must pass, and with its own defect injected into a throwaway copy, where it must
fail. The reason it exists is that a guard I wrote to catch fixed-name kernel
references used [a-z0-9_]+, which cannot match models-kernels-linux-x86_64.jar
because the name has hyphens -- it passed while the defect was present, and I
nearly committed it.

It paid for itself immediately by finding three more defects in guards that looked
green:

derive.py trusted its committed backend-delta record whenever git could not
answer, so on a shallow clone a tampered record would admit a pair whose arms
differed on the measured path; the tamper detection only worked when git did. Now
fail-closed, with an explicit env var to accept the record on trust.

With that fixed, derive.py exited 0 having rejected every pair. A derivation that
produced no numbers was reporting success. Now non-zero on zero admissible pairs
and on any rejection.

And its NOTES check passed vacuously against zero rows, so it would have blessed
any prose at all. Now non-zero when there is nothing to check against.

All three are one shape: a check reporting green because it had nothing to check.
None was visible from a clean run.

Adding a guard now means adding its mutation. If you cannot write a mutation that
makes your guard fail, the guard does not work yet.
… it working

scripts/integrity/README.md quoted the fabricated changelog path verbatim as an
example of the defect, and the check cannot tell a citation from an illustration,
so it flagged it. CI caught that, not me.

Fixed by not writing a literal non-existent path into a document at all: the
directory is named without the benchmark-results/ prefix, which reads the same and
matches nothing. The alternative was a marker comment the author has to remember,
and a guard whose correctness depends on being remembered is the kind that quietly
stops working.

The rule is now written down in that README under its own heading, because the next
person to document a bad path will hit exactly this.
…not a forgery

I had called the changelog defect "fabricated", in a mutation name and in a
pull-request body. That word is wrong and it is damaging in a permanent public
record for a commercial catalog.

Wrong, because it describes manufacturing false evidence to deceive. What happened
is that a changelog pointed at benchmark-results/2026-10-08-q4-1-exact-path when
the real directory is dated 2026-10-09-q4-1-support. The measurements were real,
taken on the harness, and every number in that entry traces to the actual NOTES.md
-- re-verified in this pass, 11 of 11. A wrong reference to real evidence is a
careless citation, not a forgery.

Damaging, because this repository exists partly to keep that distinction sharp.
Using the word for an incorrect path flattens it against the failure the house
rules were written for, and anyone reading the history later cannot tell which
one occurred.

The mutation is now evidence-references/unresolvable-path, which is what it
injects and what the guard detects. The pull-request body that used it has been
corrected to "unbacked", which is the gate's own verdict name.

One pushed commit message still carries the word. It is not being rewritten: the
branch is shared, and a correction that is visible in the history is better than
one that edits the record to look as though nothing happened.
derive.py verifies that a paired measurement's two arms came from builds whose
delta misses the measured path, which needs the commit range. On CI's shallow
clone it cannot, so once it was made fail-closed it went red on a clean tree and
the mutation check correctly refused to draw any conclusion from it.

fetch-depth: 0 rather than letting it trust a record it cannot verify. That trust
is the hole that ran a September kernel for a whole campaign with no symptom.
@bsbodden
bsbodden merged commit d85ebf0 into main Oct 10, 2026
5 checks passed
@bsbodden
bsbodden deleted the fix/fleet-kernels-jar-required branch October 10, 2026 14:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant