Repository navigation
MeshWeaver#890 recurred 3x in 30h with the leg-5 reading it was closed waiting for — reopen-or-new-path call needed #5212
Description
Activity
Two amendments to the body above — it was minted from the filing a few minutes before both were established.
1. The rate is a floor, and the floor now states what it could not see. A
cancelledshard can
carry the defect (filtering a population on its conclusion is the exact mistake MeshWeaver#890's
2026-09-12 correction records), so the 22 cancelled executions were read rather than assumed: 18 logs
clean of bothcanary=BELOW-ROSLYNandPROCESS CANNOT EMIT, 4 with no retrievable log. 3 of 96
determined, 6 undetermined.2. The "next measurement" as stated could not have been carried out, and this is the part worth
carrying forward. Both #890 and the leg's own printed residual said "repeat the private emit until it
tiers up".PrivateRoslynCopy.Emitcreates its own collectibleAssemblyLoadContextper invocation
and unloads it infinally(src/MeshWeaver.Compiler/PrivateRoslynCopy.cs:85,:113-115), so N calls
are N cold copies — the same tier-0 code measured N times, tiering nothing. The experiment needs
ONE copy held loaded while the emit is repeated inside it, with the unload deferred to the end: a
second entry point, not a loop at the existing one. Caught by the automatic review on
#5211, which corrects it in the XML doc, in the message an
operator actually reads, and inDoc/Architecture/NodeTypeCompilation.That PR also carries the occurrence table, the
flat=phase finding, and the four dead ends (platform
set,mainas a control, theexit 124arithmetic, the straggler capture). It does not attempt a fix:
the reopen-versus-new-path call on #890, and the tier-up entry point, both belong to whoever owns that
issue.A fourth occurrence, and a four-day denominator. Widening the sweep to 2026-09-19 found one more,
which also strengthens the attribution: the sets now span four, not three.occurrence run / job platform set (core) branch flat=compiler=2026-09-19 08:22Z 35430730563/1058670582253.0.0-ci.8963(31be311dc)fix/4740-plugins-runner-timeoutsEMITS×52PRIVATE-COPY-EMITS×522026-09-21 13:36Z 35606133491/1063551810933.0.0-ci.9088(88cd4b7b4)fix/incident-fold-dedupeSAME-FRAME@…×40×40 2026-09-21 22:41Z 35661949332/1065446687203.0.0-ci.9108(b1aa8a9e5)fix/hosting-bake-probeEMITS×6 thenSAME-FRAME@…×52×58 2026-09-22 11:37Z 35722501930/1067425550613.0.0-ci.9168(620a4893a)fix/ai-stream-cancelEMITS×58×58 Denominator. 201 executions of
portal-hosts (Hosting.Monolith.Test)since 2026-09-19T00:00Z: 153
success, 39cancelled, 1 running, 8failure— 4 of the 8 are this defect; the other 4 carry no
canary=and noexit 124at all. 🚨 All 201 arepull_requestruns. The unit was not selected on
mainonce in four days, so the absence of amaincontrol is a zero denominator over the wider window
too, not just the narrow one.Core's own CI over the same days is a measured negative, and a weak one: 24
MeshWeaver Build and Testruns since 2026-09-20, 13 failed jobs between them, none carryingPROCESS CANNOT EMIT,
canary=BELOW-ROSLYNorexit 124. That localises the observation to the Plugins Monolith unit; it
does not exonerate core, because 13 failures is thin and core's suites do not drive in-process NodeType
compilation at that volume.And the fault is ACQUIRED, then progressive — both measured. In the newest occurrence
OverlaySelfHealInstanceRecycleTest.OverlaidInstance_SelfRecycles_WhenTypeCompilesGreenpassed at
12:52:07.654Z, and that test asserts a NodeType reachingCompilationStatus.Okwith a usable build
whose instance then renders a marker — 33 seconds before the first poisoned emit at 12:52:40Z. So
nothing arrived broken. The 22:41Z occurrence supplies the other half inside one process:flat=EMITS
for six failures, thenflat=SAME-FRAMEfor fifty-two. That is the shape a tier-up produces and not
the shape a bad file produces, which is what makes the (corrected) tier-up experiment worth running.All of it is in #5211.
The cleanest control arrived on its own, and it is worth recording as the canonical one for this
failure mode.Attempt 2 of run
35722501930re-ran the same job (106780475056, 14:07:57Z) with head
353b8137bunchanged andplatform mount: set 3.0.0-ci.9168 (core 620a4893a)unchanged, on the
same runner pool, and reportedTest run summary: Passed! total: 813, failed: 0with zero
canary=BELOW-ROSLYN, in 14 minutes against the 900 s cap. Attempt 1 of that same run had 58 canary
blocks and died atexit 124.One comparison excludes the pull request's diff, the platform set and the content of the suite at
once — none of them changed between the attempts.🚨 And it names why this defect is expensive to read. A re-run clears it. To whoever re-runs, that
reads as a flake; to whoever only saw attempt 1, it reads as the pull request's fault. It is neither —
it is a process that stopped being able to emit, and the next process that did not. The instrument that
tells the difference is the emit canary in the job log: a control compilation failing alongside the
real ones. Without that line, "some in-mesh compiles NRE'd" reads as bad content in those NodeTypes.Recorded in #5211.
- addedsev:MMedium - secondary path broken, or easy workaroundMedium - secondary path broken, or easy workaround
on Sep 23, 2026 Classification: same defect as #890, NOT a new path — and #890 was never closed on a fix, so this is not a regression either. Keep THIS issue as the live tracker; leave #890 closed as the instrument history.
Why, from #890's own record:
- CI: one Roslyn Emit poisons the whole process — every later compile NREs in Cci.MetadataWriter.GetConsolidatedTypeParameters, shard ends exit=124 #890 was closed 2026-09-13T09:47Z (
completed) when leg 5 of the canary — the pristine-compiler control, feat: leg 5 of the #890 canary — the pristine-COMPILER control (compiler=PRIVATE-COPY-EMITS / THREW) #4165 — was built. Its closing condition, stated on 2026-09-13, was "the reading on an occurrence": an instrument, not a remedy. No change to Roslyn's use, the reference set or the emit path was merged against it after that. - All four occurrences here carry the same signature CI: one Roslyn Emit poisons the whole process — every later compile NREs in Cci.MetadataWriter.GetConsolidatedTypeParameters, shard ends exit=124 #890 was about (
PROCESS CANNOT EMIT (#890),canary=BELOW-ROSLYN,dissect=READS-HEALTHY,exit 124onportal-hosts (Hosting.Monolith.Test)), and the leg-5 reading CI: one Roslyn Emit poisons the whole process — every later compile NREs in Cci.MetadataWriter.GetConsolidatedTypeParameters, shard ends exit=124 #890 was waiting for (compiler=PRIVATE-COPY-EMITS) on every one. A "different path sharing the fingerprint" would need a reading that disagrees somewhere; none does. - So the reopen-vs-new question resolves to: the defect is unfixed and live; its diagnosis advanced (BELOW-ROSLYN void; the fault is in the shared Roslyn copy's image, its mapping, or the native code tiering produces for it;
flat=is a phase). Reopening CI: one Roslyn Emit poisons the whole process — every later compile NREs in Cci.MetadataWriter.GetConsolidatedTypeParameters, shard ends exit=124 #890 would bury that under 40 days of superseded readings; this issue carries the current state and fix(#890): the first fourcompiler=readings — and the tier-up instruction they made testable, which was wrong #5211 committed it toDoc/Architecture/NodeTypeCompilation.
The next step, precisely (from the 2026-09-22 amendment — the old instruction could not be executed): a second entry point that holds ONE
PrivateRoslynCopyloaded and repeats the emit inside it until it tiers up, deferring the unload to the end (PrivateRoslynCopy.Emittoday creates and unloads a fresh collectible context per call,src/MeshWeaver.Compiler/PrivateRoslynCopy.cs:85,:113-115, so N calls are N cold copies). If the held copy starts failing after tier-up, the fault is in what tiering produces; if it never does, it is the shared image or its mapping. This is diagnosis work on the owner's instrument, not a fix I can make from the evidence here.Not established: the rate since 2026-09-22 14:30Z (I did not re-sweep
portal-hostsruns); whether a corrupted static reachable only from theMetadataWritercall site is excluded (the 2026-09-22 body already says it is not).- CI: one Roslyn Emit poisons the whole process — every later compile NREs in Cci.MetadataWriter.GetConsolidatedTypeParameters, shard ends exit=124 #890 was closed 2026-09-13T09:47Z (
It now happens on
main: two occurrences today in a new suite, both starting at the same test, about 6 % of runs of that leg.Plugins run / job (attempt 1) set (core) canary / dissect / flat / compiler first poisoned compile 35974400848 / 107555406130 3.0.0-ci.9287 (bf5ac85) BELOW-ROSLYN ×43 / READS-HEALTHY / EMITS / PRIVATE-COPY-EMITS Widget/Thing08:45:39Z35985209459 / 107589527570 3.0.0-ci.9296 (8d3fd9f) ×41, same readings Widget/Thing10:27:55ZBoth are
portal-hosts (network-133 · leg 4/8)runningMeshWeaver.PluginCatalog.Teston runtime 10.0.12 linux-x64 with Roslyn 5.9.0, and both went green on attempt 2. The stack matches every earlier occurrence:CreateIndicesForNonTypeMembers→getConsolidatedTypeParameters×2 →NamedTypeSymbol.ITypeDefinitionMember.get_ContainingTypeDefinition.- Denominator: 2 of 33 executions of that leg since 09-23T00:00Z. The two newest main runs (36002604963, 35997990577) are red for other reasons and carry no canary.
- Onset position: identical in both. It is the same test start in the leg's deterministic order,
RetiredNodePruneTest.SharedSourceChange_…compilingWidget/Thing. Earlier emits in the same host had succeeded (ADependentsSource_CompilesAgainstTheTypeItsDependencyShipsassertsOkand passed about 2 minutes before). So the same workload poisons about one process in 16, which means something is decided per process. - Population harness, 0 occurrences: ≥600 short processes × 1,500 emits with the pipeline's options (osx-arm64, runtime 10.0.11), where about 36 were expected. The 800k-emit loop on 09-10 was a single process, so it was a single draw.
- Ruled out as the trigger: core fix(compiler): keep NodeType compiles off the silo's shared ThreadPool #5327/fix(compiler): bound compile CPU on a dedicated-thread IIoPool lane #5375 (the CompileCpu lane,
ConcurrentBuildoff). The defect predates them by a month. Shared references and the sharedCompilationare also out: legs 1 and 2 exclude them, and the private copy emits. - Hypothesis, not established: the tier-1 code of the guard's call path.
AsNestedTypeDefinitionImpl's tier-1 inlines_containingSymbol as NamedTypeSymbolbehind GDV plus a profile-guided cast expansion. Its shape varies per process, which fits per-process randomness, acquire-then-deepen, the cold private copy's immunity, and a consistent frame. - Instrument: Systemorph/MeshWeaver.Plugins#2357 writes
DOTNET_JitDisasmlistings of the six methods on that path intotest-logs/jit-890-listing.log. The file is uploaded only when a suite fails. The next occurrence will therefore carry the exact native code that answered the guard. The write-up and how to read the listing are in docs(#5212): #890 on main — one onset position, a per-process harness null, and the JIT-listing capture #5663 (Doc/Architecture/NodeTypeCompilation).
Not established: whether production replicas reach this state. A
memexLogsread ofPROCESS CANNOT EMITover 3 h returned 0, with no positive control. The 12 h and 3-day reads on memex and memex-cloud were refused by Loki (timeout, then 429). The harness null is arm64 only. In the 10:27 host, aTypeLoadException … format is invalidonBulkPack/Thing51 s before onset is recorded but not concluded from: that text is also the known load-during-unload tear, and the 08:45 host does not show it.Severity is a triage call I'm leaving open. At
sev:Mthis now reddensmainin about 6 % of runs of this leg.Census correction to my comment above: the full sweep has now finished. It covers 87 executions of
portal-hosts (network-133 · leg 4/8)from 09-23T00:00Z to 09-24T14:00Z, and 84 of them ran PluginCatalog.Test. 3 carry the canary (≈3.5 %), not 2 of 33. The third is run 35906530904 / job 107342449379 on 09-23 at 19:34Z, set 3.0.0-ci.9263 (core 9a61fed), readingsBELOW-ROSLYN/READS-HEALTHY/flat=SAME-FRAME/PRIVATE-COPY-EMITS×43. Its first poisoned compile is againWidget/ThinginRetiredNodePruneTest.SharedSourceChange_…, so all three occurrences start at the same onset test. #5663 carries the corrected table.- added a commit that references this issue
on Sep 24, 2026 Workflow-wide sweep finished. I read every non-success job of every Plugin Catalog CI attempt from 09-23T06:46Z to 09-24T14:00Z (1,060 job logs). It finds six occurrences in 31 h, four of them on main: the three in
network-133 · leg 4/8(PluginCatalog.Test, all starting atWidget/Thing), plus three innetwork-133 · leg 1/8(Hosting.Monolith.Test). Those are 35828369047 (main,TestData/UploadLegWedgeType), 35846873901 and 35858189010 (PRs, both starting atTestData/SatisfiedType). #5663 carries the list.New attributable occurrence blocking the 9945 fleet arm: Plugins core-candidate run 37211768667, leg 1/12 job 111464402311 (Hosting.Monolith.Test), core
87422043+ Pluginsfd3dbd1e. The host emitted successfully before the first poisoned compile at 2026-10-04 15:15:13 UTC (Systemorph/LinkedInProfile); thenPROCESS CANNOT EMIT (#890)recurred across unrelated NodeTypes. Its first canary has shared/pristineNullReferenceExceptionatNamedTypeSymbol.Microsoft.Cci.ITypeDefinitionMember.get_ContainingTypeDefinition,dissect=READS-HEALTHY,flat=EMITS,compiler=PRIVATE-COPY-EMITS. The leg kept running through many failing assertions, was canceled at 15:53:29, and uploaded a log but no.trx. The paired verdict correctly reportsmissingEvidence: 1, soarm-promoted-setrefused 9945 at 16:15:12; production remains 9939. This is the established acquired process-wide emit failure, not evidence that the 9945 diff caused it. I am investigating the missing JIT capture/root cause; the gate must stay red until there is a valid verdict.Candidate core-pair run 37211768667 exposed an evidence gap: the ordinary Plugins portal-host lane captures the targeted Roslyn JIT listing, but the core-candidate leg did not, and this canceled host left no .trx. Plugins PR #2873 now adds the same JIT capture to candidate/control leg artifacts and documents it. Local actionlint and workflow-shell parsing pass. The candidate verdict remains red on missing test evidence. This is instrumentation for the next attributable occurrence, not a root-cause fix or a reason to arm 9945; the PR itself is waiting on the required review/CI gate.
4 remaining items
Investigation 2026-10-06/07: two findings and one instrument (#6220). No root cause, no fix.
1. The mapping CAN change mid-process, so leg 5's remaining candidates were not "native code only". An IL-only assembly loaded by path runs out of a live file mapping. An in-place rewrite of its file (truncate + write, which is what
File.WriteAllBytesandFile.Copy(overwrite:true)do, or any write through a hard link to the same inode) is read by the SAME process. Both reflection and the IL the JIT will compile come from the new bytes.- Measured in a scratch console (osx-arm64): after the rewrite,
GetMethodreturned null in the same process. - Measured again as a committed deterministic test,
RoslynImageIntegrityTest.AnAssemblyRewrittenInPlace_IsReadFromTheNewBytes_InTheSameProcessin diag(#5212): canary leg 6 reads the compiler image; probe for the 'format is invalid' shape #6220. - That reproduces CI: one Roslyn Emit poisons the whole process — every later compile NREs in Cci.MetadataWriter.GetConsolidatedTypeParameters, shard ends exit=124 #890's shape: the fault is acquired mid-process, deepens as code tiers up, and a private copy loaded from fresh bytes is immune.
- I found NO such writer in core
src/testor in Plugins: no in-place write to anyMicrosoft.CodeAnalysis*.dllpath. So this is a candidate the next occurrence can now confirm or exclude. It is not a finding.
2. Shape b (
format is invalid) is the same event: same onset, same failing tests. In run 37535577344 (job 112524799788,format is invalid) and run 37532737494 (job 112514066537,canary=BELOW-ROSLYN):- The onset position is identical in the test order: the first nested-type emit after
NodeTypeOnDemandAdoptionTest(LinkedInProfile_NodeType_CompilesAndRendersOverview), at 22:19:18.386Z and 21:44:55.676Z respectively. - Same platform set (3.0.0-ci.10095), same runner scale set (
aks-silos-dind-trunk-q2frb). Main's green leg-1 job 112515268219 ran on that scale set too. - The emit canary only runs when Emit throws, so shape b had no verdict.
3. Things I checked and excluded:
- The CI Roslyn is the stock NuGet IL image. The leg's
Microsoft.CodeAnalysis.CSharp.dll(fc571a66…) andMicrosoft.CodeAnalysis.dll(55618d45…) from the portal-hosts-bin artifact are byte-identical to the 5.9.0lib/net10.0files. Not ReadyToRun, not a different build. - Core occurrences ran on GitHub-hosted
ubuntu-latest(jobs 112147039747 and 111891768728), so this is not specific to the AKS runner nodes.
4. Local repro attempts, all null:
- osx-arm64 (runtime 10.0.11), 4 threads: a synthetic workload (nested generics, lambdas, async, iterators, records, anonymous types), interleaved with the exact broken
config => this is not valid C# at all ((await (compile and nested and flat canaries, each image loaded and its types realised. 0 failures. - linux-x64 runtime 10.0.12 under qemu (
DOTNET_EnableWriteXorExecute=0, which CI does not set): 30 processes × 300 iterations × 4 threads. 0 failures. - So the synthetic workload does not reach the state on either architecture. That is the same null this issue has had all along.
5. Instrument shipped in #6220 (diagnostic only; no verdict, status or classifier change):
image=(leg 6). One baseline per process at the first emit: the mapped metadata viaTryGetRawMetadata, the IL of about 11k guard-path methods read through the runtime, and the file's SHA-256. Each failure is then classified asINTACT,REWRITTEN-IN-PLACE,MAPPED-CHANGED,FILE-REPLACEDorNO-BASELINE.image=INTACTtogether withcompiler=PRIVATE-COPY-EMITSleaves the shared copy's native code as the only remaining candidate: a JIT tier-1/PGO miscompile or code-heap corruption.- Related open JIT bugs on 10.0.12 exist, for example JIT: (bug) Failed
isinstto an interface is treated as non-null, folding a GetType() compare and swallowing a NullReferenceException dotnet/runtime#134193 (a failedisinsttreated as non-null). That one is interface-only, so it is not this site.
host=records CPU model, avx2/avx512f and any tiering knobs, which no occurrence recorded before.loadcanary=covers shape b: on aformat is invalidload failure it emits a known-good nested canary through the shared compiler and loads it.
6. Incidental: the #890 JIT listing has never been captured. Neither 2026-10-06 failure artifact contains
jit-890-listing.log, although the per-test logs in the sametest-logs/directory are present.- The same
DOTNET_JitDisasm+DOTNET_JitStdOutFilesettings DO write the file locally, on both osx-arm64 and linux-x64 10.0.12 release runtimes. Tier1 listings ofAsNestedTypeDefinitionImplandgetConsolidatedTypeParametersappear. - So the env var is not reaching the MTP test host in the Plugins
portal-hosts-run-suites.shpath, or the file is lost. - Fixing that capture (Plugins) is the other half of what the next occurrence needs.
Not established:
- The root cause.
- Whether any process in CI rewrites a Roslyn image.
- Whether the native-code hypothesis holds. Nobody has run the split-arm
DOTNET_TieredPGO=0experiment yet.
- Measured in a scratch console (osx-arm64): after the rewrite,
Another occurrence: core PR #6231, head
64317ea803, run job 112723124834 (shard 3,Memex.Portal.Shared.Test).- Failure:
NullReferenceExceptionatNamedTypeSymbol.Microsoft.Cci.ITypeDefinitionMember.get_ContainingTypeDefinitioninMetadataWriter.PopulateNestedClassTableRows. The compile watcher loggedPROCESS CANNOT EMIT (#890)withcanary=BELOW-ROSLYN. - First poisoned compile:
InstallNeverInheritsRepoCompileVerdictTestat 09:30:42Z, the 69th of 150 test classes in that process. After it,ModulesUpdateIndependentlyOfThePlatformTest(2 tests) failed with the same NRE. - Not the PR: the PR doesn't touch the compiler, and the same classes pass locally on the same head (7/7). Shard 3 passed on the branch's previous head.
- Not established: what poisoned the process before 09:30:42Z, because successful compiles aren't logged.
- Failure:
PR #6231 shard 3 (2 of 2): what a local reproduction attempt established. It did NOT reproduce.
Scope. I set out to reproduce the 2-of-2 onset on PR #6231 locally, bisect the preceding classes, find the mechanism and fix it. None of the 11 local processes reproduced it, so there is no minimal repro and no fix in this comment. Below is what the attempt established and what it excluded. Some of the findings correct earlier statements on this issue.
1. The two CI processes
- Jobs
112723124834and112729207042ran the same tested tree: head64317ea80merged ontomain@fa6bb901bf(fresh merge6930b34bcf38). The tests ranMemex.Portal.Shared.Testpart 1/2 (161 of 321 classes, every second class from the sorted list), with seed1319993060,parallel mode = none, on .NET 10.0.12 linux-x64. - The method order is identical in both jobs (280 test methods, diffed). The first poisoned compile is
VerdictPkg/CataloginInstallNeverInheritsRepoCompileVerdictTest, the 70th class. The 69 classes before it areSettingsRowActionsActOnTheClickedRowTest(Fix chat composer losing focus when prompt ends with question mark #69) and earlier; the full list is reproducible from either log. - First canary reading, job 1, 09:30:42.611Z:
canary=BELOW-ROSLYN, shared and pristine both THREW atNamedTypeSymbol…get_ContainingTypeDefinition,dissect=READS-HEALTHY,flat=SAME-FRAME,compiler=PRIVATE-COPY-EMITS,held=EMITS(64/64),image=INTACT(metadata, IL of in-scope methods and file hashes all unchanged),host=(.NET 10.0.12 linux-x64 AMD EPYC 7763, 4 cores, avx2, no avx512, tiering=defaults). - Between the last test that is known to have compiled successfully (
DiagnosticsNameTheContentTypeAssemblyTest, 09:30:17) and the onset, neither log contains a TypeLoad, BadImageFormat, "format is invalid", unload or AccessViolation line. - The onset position is not tied to one preceding class. Across the three core occurrences the onset test differs:
LinkVerdictPerPlatformTest(2.4 #57, job111891768728),ModulesUpdateIndependentlyOfThePlatformTest(job112147039747), andInstallNeverInherits…(Fix incorrect chat composer behavior when prompt contains slash character #70, today).SettingsRowActions…andPackageUninstallTestcome before the onset in two of the three, but not in the 10-05 job. What all three share is the position: the first compile about 4 to 5 minutes into the process, after a stretch of classes that do not compile.
2. 🚨 Correction: CI's seed does not reproduce CI's order locally
xUnit v3 orders test collections by unique IDs that hash the assembly path, so
:<seed>with the same class set runs in a different order on another machine. Measured: same seed, same 161 classes, and locally the first class wasStandInCompileFailureNamesTheEmitCanaryTestwhere CI's wasAFailedPackageNamesItsCauseTest. The 10-06 "full run with CI's seed, 0 failed" therefore did not run CI's order. To replay CI's exact order I used a local-only[assembly: TestCollectionOrderer]plus[assembly: TestMethodOrderer]that read the class and method order from the CI log. The resulting method order diffed empty against CI's.3. Local attempts (all on the merged tree, built Release; 0 of 11 processes poisoned)
environment runs order result macOS arm64, .NET 10.0.11 1 CI class set, local order 1305 passed, no canary linux-arm64 container, 10.0.12, 4 CPUs 1 exact CI order 380 passed, no canary macOS x64 (Rosetta 2), 10.0.12 5 exact CI order (3 of the 5 with DOTNET_PROCESSOR_COUNT=4)all clean macOS x64 (Rosetta), TC_CallCountThreshold=5,TC_CallCountingDelayMs=01 exact CI order clean macOS x64 (Rosetta), full suite (321 classes) 1 default 2409 tests, 1 unrelated failure (no x64 SDK for dotnet msbuild), no canarylinux-x64 under qemu ( EnableWriteXorExecute=0, which qemu needs)1 exact CI order 380 passed, no canary Two more harnesses also came back clean:
- Standalone emit population on x64 (Rosetta): 160 processes × 300 emits. Each process uses a randomized mix of nested and generic types, lambdas (display classes), static lambdas, async and iterator state machines, anonymous types, records and LINQ, emitted with portable PDB and XML docs, with the flat canary after every emit. 0 poisoned. All four writer methods reach Tier1 in this harness.
- Roslyn-copy probe (a local class placed just before the onset class): there is exactly one
Microsoft.CodeAnalysis/.CSharpin the process, in the Default ALC, and no collectible ALC is alive at that point.
4. What
DOTNET_JitDisasmSummaryshows on x64 (new)- Under CI's exact order with default tiering, the emit path never leaves Tier0 locally. In 5 of 5 runs,
MetadataWriter.PopulateNestedClassTableRows,getConsolidatedTypeParameters,FullMetadataWriter.CreateIndicesForNonTypeMembersandNamedTypeSymbol.AsNestedTypeDefinitionImplwere still at (Instrumented) Tier0 through class 71. So locally, class 70 runs unoptimized emit code. Whether CI's process had promoted these methods by then is unknown; nothing in a CI log records it. - When they do reach Tier1 locally (with the lowered threshold, in the full suite, and in the harness), every one is
Tier1 with Synthesized PGO, and none poisons. The Tier1 listing ofAsNestedTypeDefinitionImplon x64 is correct. It does a GDV onSourceNamedTypeSymboland inlines_containingSymbol as NamedTypeSymbolasnull → null,MT == SourceNamedTypeSymbol → self, otherwiseCORINFO_HELP_ISINSTANCEOFCLASS. That is the right null polarity for a top-level type. - This sharpens the codegen hypothesis without confirming it. If the cause is Tier1 code, it needs a Dynamic-PGO profile shape that none of the local runs produced. The onset falling on the first compile after a quiet stretch fits that: call counting and tier-up happen in idle windows, such as the four 500 ms
NotEmitwaits inSettingsRowActions…. That is a fit, not proof.
5. Excluded by reading the runtime source (release/10.0)
- A stale collectible type in a PGO profile.
JIT_ClassProfile32/64recordsDEFAULT_UNKNOWN_HANDLEfor collectible types, and the delegate and vtable probes skip collectible MethodDescs (jithelpers.cpp1957–2189). - A stale CastCache entry after an unload.
LoaderAllocatorflushes it (CastCache::FlushCurrentCache(),loaderallocator.cpp:631). - A shared System.Reflection.Metadata or System.Collections.Immutable.
PrivateRoslynCopyloads onlyMicrosoft.CodeAnalysisand.CSharpprivately and shares those two from Default, yet it emits. That clears them. - Corrupted pooled objects or other Roslyn static state do not explain
flat=SAME-FRAME.PopulateNestedClassTableRowsand the guard path use no pooled objects, and a member-less top-level class can only reach the frame through a guard misread. A corrupted static that is reachable only from the writer's call site is still not formally excluded.
6. Incidental finding, found along the way
A stock Linux container (
fs.inotify.max_user_instances=128) fails 305 of 380 tests in this shard withIOException: The configured user limit (128) on the number of inotify instances has been reached. The trace runsServiceSetup.CreateServiceCollection(test/MeshWeaver.Fixture/ServiceSetup.cs:37) →JsonConfigurationSource.Build→PhysicalFilesWatcher. That looks like one config file watcher per test that is never released. GitHub runners evidently allow more instances, so CI does not see it. Not triaged further.7. Next step
This composition reproduced 2 of 2 on CI, so one CI run of shard 3 on this tree is decisive in a way 11 local processes were not. Run it with two arms:
- (a)
DOTNET_JitStdOutFile+DOTNET_JitDisasmSummary=1+DOTNET_JitDisasm="PopulateNestedClassTableRows AsNestedTypeDefinitionImpl *getConsolidatedTypeParameters* CreateIndicesForNonTypeMembers", which shows whether a Tier1 version of these methods existed before 09:30:42 and what code it contained; - (b)
DOTNET_TieredPGO=0, the split arm.
If (b) is clean and (a) shows a Dynamic-PGO Tier1 body with an inverted null branch, the venue is dotnet/runtime, and that listing is the report. If (a) shows the methods still at Tier0 at onset, the codegen hypothesis is dead and the search moves to process state.
I did not run that experiment: this task excluded CI polling.
- Jobs
Three more core occurrences on 2026-10-07, all shard 3
Memex.Portal.Shared.Test(part 1/2, 322 class args). They were reported as "caused by #6208 / the fresh-merge base". They are not.run / job event, tested tree first bad reading failing tests 37612813368 / 112766252743 push main cbb3053100(11:15Z), before #6208 mergedcanary=BELOW-ROSLYN×11InstallNeverInherits…, ModulesUpdateIndependentlyOfThePlatformTest ×2 37637952561 / 112852969409 PR #6231 head a9fa0f8dbbfresh-merged onto maincac0664de14:49:36Z PROCESS CANNOT EMIT,canary=BELOW-ROSLYN(~4m41s into the process)the same 3 37637806820 / 112852220267 PR #6244 (workflow/doc only) fresh-merged onto main cac0664de14:52:50Z CS0246 on a stand-in passed in the reference list, canary=OK(~3m57s in)LinkVerdictPerPlatformTest ×1, ModuleLinkVersionTest ×11, then the same 3 New in the third one: the shallow phase reads
canary=OK. The first 12 failures are stand-in compiles (StandInCompile.Emit) where the extraMetadataReference.CreateFromImage(...)is invisible (CS0246 'YamlDotNet'×9,CS0246 'PerPlatformVerdictFromTheFuture',CS0234 'MeshWeaver.Test'×2) while the canary, which compiles against the same reference set WITHOUT an extra image reference, still emits. 15 s later (14:53:05Z) the same process readscanary=BELOW-ROSLYN(shared:andpristine:both THREW NRE atMetadataWriter.<GetConsolidatedTypeParameters>). So the canary misses the onset; the "invisible extra reference" symptom comes first in this process.Non-determinism, same inputs: both PR runs tested a fresh merge onto the same main
cac0664dewith nothing in their diffs that reaches the compiler, andLinkVerdictPerPlatformTest/ModuleLinkVersionTestpassed in the #6231 process and failed in the #6244 process. Main's own push run ofcac0664de(37625265553) passed shard 3.Not established: what poisoned either process before the onset (successful compiles are not logged), and I did not attempt a local repro (the 11 null attempts above make it a poor use of time without the JIT/tiering capture from
experiment/5212-jit-capture).Three more core occurrences on 2026-10-07, all shard 3
Memex.Portal.Shared.Test(part 1/2, 322 class args). They were reported as "caused by #6208 / the fresh-merge base". They are not.run / job event, tested tree first bad reading failing tests 37612813368 / 112766252743 push main cbb3053100(11:15Z), before #6208 mergedcanary=BELOW-ROSLYN×11InstallNeverInherits…, ModulesUpdateIndependentlyOfThePlatformTest ×2 37637952561 / 112852969409 PR #6231 head a9fa0f8dbbfresh-merged onto maincac0664de14:49:36Z PROCESS CANNOT EMIT,canary=BELOW-ROSLYN(~4m41s into the process)the same 3 37637806820 / 112852220267 PR #6244 (workflow/doc only) fresh-merged onto main cac0664de14:52:50Z CS0246 on a stand-in passed in the reference list, canary=OK(~3m57s in)LinkVerdictPerPlatformTest ×1, ModuleLinkVersionTest ×11, then the same 3 New in the third one: the shallow phase reads
canary=OK. The first 12 failures are stand-in compiles (StandInCompile.Emit) where the extraMetadataReference.CreateFromImage(...)is invisible (CS0246 'YamlDotNet'×9,CS0246 'PerPlatformVerdictFromTheFuture',CS0234 'MeshWeaver.Test'×2) while the canary, which compiles against the same reference set WITHOUT an extra image reference, still emits. 15 s later (14:53:05Z) the same process readscanary=BELOW-ROSLYN(shared:andpristine:both THREW NRE atMetadataWriter.<GetConsolidatedTypeParameters>). So the canary misses the onset; the "invisible extra reference" symptom comes first in this process.Non-determinism, same inputs: both PR runs tested a fresh merge onto the same main
cac0664dewith nothing in their diffs that reaches the compiler, andLinkVerdictPerPlatformTest/ModuleLinkVersionTestpassed in the #6231 process and failed in the #6244 process. Main's own push run ofcac0664de(37625265553) passed shard 3.Not established: what poisoned either process before the onset (successful compiles are not logged), and I did not attempt a local repro (the 11 null attempts above make it a poor use of time without the JIT/tiering capture from
experiment/5212-jit-capture).Three more core occurrences on 2026-10-07, all shard 3
Memex.Portal.Shared.Test(part 1/2, 322 class args). They were reported as "caused by #6208 / the fresh-merge base". They are not.run / job event, tested tree first bad reading failing tests 37612813368 / 112766252743 push main cbb3053100(11:15Z), before #6208 mergedcanary=BELOW-ROSLYN×11InstallNeverInherits…, ModulesUpdateIndependentlyOfThePlatformTest ×2 37637952561 / 112852969409 PR #6231 head a9fa0f8dbbfresh-merged onto maincac0664de14:49:36Z PROCESS CANNOT EMIT,canary=BELOW-ROSLYN(~4m41s into the process)the same 3 37637806820 / 112852220267 PR #6244 (workflow/doc only) fresh-merged onto main cac0664de14:52:50Z CS0246 on a stand-in passed in the reference list, canary=OK(~3m57s in)LinkVerdictPerPlatformTest ×1, ModuleLinkVersionTest ×11, then the same 3 New in the third one: the shallow phase reads
canary=OK. The first 12 failures are stand-in compiles (StandInCompile.Emit) where the extraMetadataReference.CreateFromImage(...)is invisible (CS0246 'YamlDotNet'×9,CS0246 'PerPlatformVerdictFromTheFuture',CS0234 'MeshWeaver.Test'×2) while the canary, which compiles against the same reference set WITHOUT an extra image reference, still emits. 15 s later (14:53:05Z) the same process readscanary=BELOW-ROSLYN(shared:andpristine:both THREW NRE atMetadataWriter.<GetConsolidatedTypeParameters>). So the canary misses the onset; the "invisible extra reference" symptom comes first in this process.Non-determinism, same inputs: both PR runs tested a fresh merge onto the same main
cac0664dewith nothing in their diffs that reaches the compiler, andLinkVerdictPerPlatformTest/ModuleLinkVersionTestpassed in the #6231 process and failed in the #6244 process. Main's own push run ofcac0664de(37625265553) passed shard 3.Not established: what poisoned either process before the onset (successful compiles are not logged), and I did not attempt a local repro (the 11 null attempts above make it a poor use of time without the JIT/tiering capture from
experiment/5212-jit-capture).Three more core occurrences on 2026-10-07, all shard 3
Memex.Portal.Shared.Test(part 1/2, 322 class args). They were reported as "caused by #6208 / the fresh-merge base". They are not.run / job event, tested tree first bad reading failing tests 37612813368 / 112766252743 push main cbb3053100(11:15Z), before #6208 mergedcanary=BELOW-ROSLYN×11InstallNeverInherits…, ModulesUpdateIndependentlyOfThePlatformTest ×2 37637952561 / 112852969409 PR #6231 head a9fa0f8dbbfresh-merged onto maincac0664de14:49:36Z PROCESS CANNOT EMIT,canary=BELOW-ROSLYN(~4m41s into the process)the same 3 37637806820 / 112852220267 PR #6244 (workflow/doc only) fresh-merged onto main cac0664de14:52:50Z CS0246 on a stand-in passed in the reference list, canary=OK(~3m57s in)LinkVerdictPerPlatformTest ×1, ModuleLinkVersionTest ×11, then the same 3 New in the third one: the shallow phase reads
canary=OK. The first 12 failures are stand-in compiles (StandInCompile.Emit) where the extraMetadataReference.CreateFromImage(...)is invisible (CS0246 'YamlDotNet'×9,CS0246 'PerPlatformVerdictFromTheFuture',CS0234 'MeshWeaver.Test'×2) while the canary, which compiles against the same reference set WITHOUT an extra image reference, still emits. 15 s later (14:53:05Z) the same process readscanary=BELOW-ROSLYN(shared:andpristine:both THREW NRE atMetadataWriter.<GetConsolidatedTypeParameters>). So the canary misses the onset; the "invisible extra reference" symptom comes first in this process.Non-determinism, same inputs: both PR runs tested a fresh merge onto the same main
cac0664dewith nothing in their diffs that reaches the compiler, andLinkVerdictPerPlatformTest/ModuleLinkVersionTestpassed in the #6231 process and failed in the #6244 process. Main's own push run ofcac0664de(37625265553) passed shard 3.Not established: what poisoned either process before the onset (successful compiles are not logged), and I did not attempt a local repro (the 11 null attempts above make it a poor use of time without the JIT/tiering capture from
experiment/5212-jit-capture).Stopgap (explicit, temporary): while this is open, a core PR whose shard 3 log carries this issue's signature (
canary=BELOW-ROSLYN/PROCESS CANNOT EMIT) is treated as a test-process (harness) death and its failed jobs are re-run ONCE — the poisoned process tells us nothing about the PR. Applied 2026-10-07 to #6244, #6231, #6248 (11 signature lines each). Root cause stays open here: experiment #6245 ran both arms (A: JIT capture, B: TieredPGO=0) clean once — no occurrence, so inconclusive — and is being re-run to catch an occurrence under capture.Experiment #6245, run 37635375205 attempt 3: the JIT capture caught an occurrence
Arm A (JIT capture, default tiering) failed with this issue's signature in job
112878675728. Arm B (DOTNET_TieredPGO=0, job112878675545) passed. I read the arm A artifactjit-capture-shard3-attempt3(jit.txt, 50 MB) together with the job log and the.trx.Capture coverage
DOTNET_JitStdOutFileappends; it does not overwrite. The one file holds 5 processes run one after another. Each starts with its ownXunitAutoGeneratedEntryPoint:Main, and the summary numbering restarts at1:. TheMemex.Portal.Shared.Testpart-1 process starts at itsMain(file line 902), so the capture covers that process from its very first compile.Timeline (log and trx timestamps; JIT events placed by their position among the test classes' own first-time Tier0 compiles)
when (UTC) event 15:46:46.6 process starts its first test class 15:51:05.40–06.52 ModuleLinkVersionTestruns 11 stand-in emits, all pass. Its last new Roslyn JITs (WithPublicSign,StrongNameKeys) are the last first-time Roslyn compiles before the onset15:51:06.5 → 15:51:15.1 (during ModuleReloadTransientDownloadTest)background tier-up batch of the metadata writer (≈400 Tier1 compiles). The first one is MetadataWriter:PopulateNestedClassTableRows() [Tier1 with Synthesized PGO]— no first-time Roslyn JIT between the batch and the onset. That fits no emit running in between, but does not prove it 15:51:21.991 onset: VerdictPkg/CataloginInstallNeverInheritsRepoCompileVerdictTest, NRE atNamedTypeSymbol…get_ContainingTypeDefinition←PopulateNestedClassTableRows;canary=BELOW-ROSLYN,flat=SAME-FRAME,compiler=PRIVATE-COPY-EMITS,held=EMITS(64/64),image=INTACT,host=(.NET 10.0.12 linux-x64 INTEL XEON PLATINUM 8573C, avx2+avx512f, tiering=defaults)after the first canary getConsolidatedTypeParametersreaches Tier1. The canary's private-copy legs show up as FullOpts compiles of the same methodsTier of each method at the onset (shared copy, Default ALC)
method at onset NamedTypeSymbol:AsNestedTypeDefinitionImplTier1 (PGO), compiled long before the onset ( fgCalledCount 190). It is correct, see belowFullMetadataWriter:CreateIndicesForNonTypeMembersTier1 (PGO), long before. Not on the faulting path MetadataWriter:PopulateNestedClassTableRowsTier1 (PGO), fgCalledCount 205, 17 inlinees with PGO data. Compiled after the last successful emit; the next emit is the onsetgetConsolidatedTypeParametersInstrumented Tier0 (Tier1 only after the onset) No
OSRcompile of any of the four. Each loop method shows twoInstrumented Tier0compiles of identical code size; I have not explained that.The Tier1
PopulateNestedClassTableRowsdropsisinst NamedTypeSymbolRoslyn 5.9.0
SourceMemberContainerTypeSymbol.ContainingTypeisldfld _containingSymbol; isinst NamedTypeSymbol; ret, where the field is aNamespaceOrTypeSymbol. The Tier1 body inlinesAsNestedTypeDefinitionImplbehind a GDV on method-table0x7FB5D6EBAA60and tests the field for null:G_M000_IG15: mov rax, 0x7FB5D6EBAA60 ; cmp qword ptr [r12], rax ; jne IG34 ; GDV on typeDef, else interface call G_M000_IG09: mov rax, gword ptr [r12+0x38] ; _containingSymbol test rax, rax jne G_M000_IG22 ; non-null NAMESPACE => "has ContainingType" G_M000_IG22: mov rcx, gword ptr [r12+0x38] ; ... ; call ContainingModule ; == [module+0xB0] (SourceModule) G_M000_IG27: mov r13, r12 ; return this => "nested" G_M000_IG19: call [NamedTypeSymbol:…get_ContainingTypeDefinition()] ; ContainingType == null => NRE
Neither loop clone has a
CORINFO_HELP_ISINSTANCEOFCLASS. For a top-level source type,_containingSymbolis its namespace, which is non-null. So the guard reads TRUE, the type is written as nested, andContainingTypeDefinitionthrows. That is exactlyflat=SAME-FRAME: a member-less top-level class is enough.Control in the same process, on the same field and the same guessed class. The Tier1
getConsolidatedTypeParameters, compiled just after the onset, carries the same dropped cast in its inlined guard (IG39: mov rdi,[rbx+0x38]; test rdi,rdi; jne IG42). In the same method, another[x+0x38] as NamedTypeSymbolis expanded correctly:IG49: cmp [rax],0x7FB5D6EBAA60; je …/IG50: call CORINFO_HELP_ISINSTANCEOFCLASS. That second site explains the earlier occurrences that failed in thegetConsolidatedTypeParametersframe, and theflat=EMITS→SAME-FRAMEdeepening: one method, then the next, tiers into the same defect. The standalone Tier1AsNestedTypeDefinitionImplguessesSynthesizedReadOnlyListEnumeratorTypeSymbol. For that class the container field is typed as a sealedNamedTypeSymbolsubclass, so leaving the cast out there is legitimate.The leg readings now have an explanation. The private copy lives in a collectible ALC. Collectible code is not tiered: every compile in it is
FullOpts, with no PGO and no GDV. That is whycompiler=PRIVATE-COPY-EMITSandheld=EMITS(64/64)hold every time. It also means the held-copy experiment cannot test tier-up as designed.Verdict
This points to the JIT: Tier1 code with Dynamic PGO that inlines a GDV-guarded callee and loses a class
isinst. It matches the mechanism stated in dotnet/runtime#134193 (anisinstresult inherits the input's non-null-ness), but that issue is about interfaces, so whether it is the same defect is for the runtime team. A report draft is ready, with the listing excerpts, the in-method control and the evidence chain. It has not been filed.Proposed stopgap (not implemented)
Temporary stopgap: set
DOTNET_TieredPGO=0on the test shards. The root cause is the JIT miscompile above. The real fix is the runtime fix, or a runtime version that has it; this setting only avoids the trigger. Production portals compile NodeTypes on the same runtime and the Plugins occurrences show the same signature, so the portals are exposed too. That needs its own decision; this proposal covers test shards only.Is "PGO off prevents it" statistically meaningful? Not yet.
Arm B has 1 clean run against this capture's 1 failure. Arm A failed in 1 of 3 attempts, and recent shard-3 runs fail about 30–50 % of the time. One clean arm proves nothing on its own. I recommend ≥10 paired runs (both arms, same commit and window):
- If arm B is 0/10 and arm A is ≥4/10, Fisher one-sided p ≈ 0.04 (≈0.016 at 5/10).
- Against a base rate of ≥0.3 alone, arm B 0/13 gives p < 0.01.
The disassembly already names the mechanism, so the arms are now confirmation, not discovery.
Not established
- The method-table → class mapping is inferred, not observed. I took Roslyn 5.9.0 field offsets from a local runtime (osx-arm64 10.0.11).
SourceNamedTypeSymbol._containingSymbolis at +0x38.PEModuleBuilder.SourceModuleis at +0xB0, which matches the listing's[module+0xB0]compare. The release JIT's listing does not print class names. - That this Tier1 body was the code running at 15:51:21.991. It was compiled before the onset emit and after the last successful one, and nothing in between shows Roslyn activity. Code activation itself is not logged.
- Why this process produced that profile or inline shape and others do not. The local runs never did.
- No JitDump or checked-JIT confirmation. No local repro.
Filed upstream as dotnet/runtime#135350 (public, maintainer-approved). Searched dotnet/runtime first: no duplicate; dotnet/runtime#134193 (interface
isinstinheriting the input's non-null-ness) is cited there as possibly related, not asserted as the same defect. The public body is the draft from the analysis comment, scrubbed of job ids, repo names and artifact names, with the TieredPGO=0 evidence stated as 1 clean run vs 1 failure against a ~30-50% base rate, and an explicit 'What is NOT established' section (no standalone repro, activation not logged, MT->class mapping inferred).
MeshWeaver#890 (CI: one Roslyn Emit poisons the whole process; shard ends exit=124) is CLOSED since 2026-09-13. It has now occurred three times inside thirty hours, and each occurrence carries the exact reading that issue said would close a question. The reopen-versus-new-path decision belongs to whoever owns the fix, which is why this is a triage filing and not a reopen.
OCCURRENCES. All three are MeshWeaver.Plugins, workflow Plugin Catalog CI, job portal-hosts (Hosting.Monolith.Test), all exit 124 with no verdict:
WHAT THE READINGS SETTLE. compiler=PRIVATE-COPY-EMITS means a second Roslyn, loaded from fresh bytes into its own collectible context and never executed in this process before, emits the very source the shared copy cannot, in the same process on the same CLR and heap. So BELOW-ROSLYN is void on all three and the dotnet/runtime venue is wrong. And flat= turns out to be a PHASE, not a property of an occurrence: occurrence 2 read both values in one process, shallow first. That closes the 2026-09-08 vs 2026-09-10 contradiction recorded on the issue (one defect at two depths) and retires confined-to-GetConsolidatedTypeParameters-recursion as a general claim. Of the three candidates leg 5 leaves open (the image, its mapping, the native code produced for it), only the third can get worse while a process runs.
NEXT MEASUREMENT, as the issue already prescribes: repeat the private emit until it tiers up. If it then fails, the fault is in what tiering produces; if it never does, it is the shared image or its mapping. This does NOT yet exclude a corrupted static reachable only from the MetadataWriter call site.
WHAT IS NOT THE CAUSE. Not a pull request: three unrelated branches. Not the platform ceiling: three different sets, and the diff between core 6b3fda2 (a set that ran the same suite green, 813 of 813, run 35700505749) and core 620a489 touches nothing under src/MeshWeaver.Compiler or src/MeshWeaver.Compiler.Pipeline, with Microsoft.CodeAnalysis.CSharp 5.9.0 and global.json unchanged.
RATE WITH ITS DENOMINATOR. Since 2026-09-21T00:00Z the unit ran 102 times across the workflow: 75 success, 22 cancelled (superseded pushes), 2 running, 3 failure, and all three failures are this defect. All 102 are pull_request runs: Hosting.Monolith.Test was not selected on main once in that window (main ran Kernel.Test, network-133, Json.Test and others instead), so it passes on main is a zero denominator, not a green, and a single re-run cannot supply the control either.
USER IMPACT. Each occurrence kills an otherwise-healthy pull request run with no verdict, and the red reads as the PR's own until someone spends an hour proving otherwise.
EVIDENCE AND WRITE-UP. Issue comment: github.com//issues/890#issuecomment-5777968513. Committed to Doc/Architecture/NodeTypeCompilation in pull request github.com//pull/5211. Attribution note on the affected PR: github.com/Systemorph/MeshWeaver.Plugins/pull/2231#issuecomment-5777974872.
WHAT I DID NOT ESTABLISH. Whether this is a regression of the fix that closed 890 or a second path to the same symptom; whether a successful emit preceded the first poisoned one in occurrence 3 (successful compiles are not logged, so the evidence is indirect); and the mapping-versus-native-code split, which the tier-up repeat is the experiment for.
Origin: https://memex.systemorph.com/rbuergi/Feedback/890-roslyn-emit-poison-three-occurrences-2026-09-22
filed by the triage agent on behalf of rbuergi via systemorph-com