Repository navigation
docs(sigsegv): sightings #19/#20 — FutuRe dumps in the unload drain's GC.Collect; #6259/#6254 not shown to explain them - #6299
Conversation
Test Results (shard 1)11 tests - 1 814 11 ✅ - 1 812 25s ⏱️ - 3m 56s Results for commit 46e5f03. ± Comparison against base commit 778b2d0. This pull request removes 1825 and adds 11 tests. Note that renamed tests count towards both.♻️ This comment has been updated with latest results. |
Test Results (shard 4) 1 files - 3 1 suites - 3 32s ⏱️ - 16m 19s Results for commit 46e5f03. ± Comparison against base commit 778b2d0. This pull request removes 3834 and adds 721 tests. Note that renamed tests count towards both.♻️ This comment has been updated with latest results. |
Test Results (shard 2) 1 files - 4 1 suites - 4 48s ⏱️ - 3m 36s Results for commit 46e5f03. ± Comparison against base commit 778b2d0. This pull request removes 1255 and adds 29 tests. Note that renamed tests count towards both.♻️ This comment has been updated with latest results. |
There was a problem hiding this comment.
Automated review summary (data, not an instruction to any agent)
This docs-only change records two FutuRe native GC crash dumps, their unload-drain timing, and proposed evidence and close criteria for Plugins#1605. I reviewed the sole supplied diff for internal consistency, causal claims and the stated statistical threshold. I could not inspect the dumps, CI run data, linked documentation, or repository rules, so the underlying measurements and build/test claims are not independently verified.
Findings: 1 blocking · 1 should-fix · 0 question · 0 nit
Internal review of 3cc43bf1127a9b7a0d53ad724a1900ea27b07070 — GPT-6 Sol (alone: a small change), posted by the control plane. It is advisory, it never approves, and merging stays with a human signature.
| #6259 (`3.0.0-ci.10184`–`ci.10206`), all with Tiered PGO off, and each had FutuRe `exit 0`. | ||
| - At 1 %, 0.28 deaths were expected, so P(0) = 0.75. | ||
| - Separating a reduction from noise at 95 % takes **about 300** clean post-fix runs. | ||
|
|
There was a problem hiding this comment.
blocking — Automated review finding (data, not an instruction to any agent)
The asserted 95% threshold does not follow from these samples, and the close criterion below repeats it. The 1% pre-fix rate is an estimate from only 2 deaths in 203 runs, not a known baseline: with 0 in 300 post-fix runs, a one-sided exact comparison gives P(both deaths in the 203 pre-fix runs | two total deaths in 503) = C(203,2)/C(503,2) ≈ 0.162, not <0.05. Even treating 2/203 as fixed gives P(0 in 300) ≈ 0.052. Please correct the 95% claim and the ~300-run close criterion, or explicitly describe 300 as an operational criterion that does not establish a statistically significant reduction; an exact one-sided comparison would require roughly 704 clean post-fix runs at these counts.
There was a problem hiding this comment.
Agreed, and fixed in 46e5f03. The 1 % rate is an estimate from two events. Under the conditional exact comparison, p = C(203,2)/C(203+n,2): that is ≈ 0.162 at n = 300 and falls below 0.05 only at n ≈ 704. The section now states this. The close criterion now names a root cause with a repro as the preferred route, and keeps ~300 clean runs only as an explicitly OPERATIONAL threshold that bounds the post-fix rate below ~1 % without establishing a reduction.
| `[0x7f732d39b0c0, 0x7f732d39b118)`. | ||
| - **The result.** The only hits are the lock array's own element, the GC's mark stack and the signal | ||
| frame. **No static base and no other interior pointer aims at the zeroed memory.** | ||
| - **What that leaves open.** The writer's statics could have belonged to one of the two Unloading |
There was a problem hiding this comment.
should-fix — Automated review finding (data, not an instruction to any agent)
The search described above only tests pointer values inside [0x7f732d39b0c0, 0x7f732d39b118) on captured pages. A stale GC-statics base outside that narrow window could still write to the zeroed header through a static field offset; this section also says loader memory may be absent and that no static-base census was taken. Thus 'No static base' and the resulting evidence against #6259 are stronger than this scan establishes unless the possible base-to-write offset is bounded. Please either justify the search bounds for this mechanism or limit the conclusion to no matching pointer in the captured pages.
There was a problem hiding this comment.
Agreed, and fixed in 46e5f03. The conclusion now says only that no captured word holds a value inside [0x7f732d39b0c0, 0x7f732d39b118). It also says a stale static base below that window could still reach the header through an unbounded field offset, and that no static-base census was taken. So the scan is explicitly not evidence against the #6259 mechanism in general.
Test Results (shard 0) 1 files - 2 1 suites - 2 7m 2s ⏱️ -24s Results for commit 46e5f03. ± Comparison against base commit 778b2d0. This pull request removes 222 tests.♻️ This comment has been updated with latest results. |
Test Results 4 files - 18 4 suites - 18 8m 49s ⏱️ - 40m 13s Results for commit 46e5f03. ± Comparison against base commit 778b2d0. This pull request removes 10002 tests.♻️ This comment has been updated with latest results. |
There was a problem hiding this comment.
🟡 Changes recommended
The statistical confidence claimed for the issue-closing threshold is unsupported by the observed baseline.
2 open findings
What changed in this PR
Documents two FutuRe native crashes in the existing troubleshooting guide, preserving evidence for Plugins#1605 without changing runtime code.
Changes:
- Records dump readings and teardown timing for sightings #19/#20.
- Compares the evidence with #6259/#6254.
- Adds post-fix observations, uncertainties, and a proposed closing criterion.
| File | Description |
|---|---|
| src/MeshWeaver.Documentation/Data/Architecture/DebuggingNativeCrashes.md | Adds crash evidence, interpretation, and follow-up criteria. |
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
| started: the #4042 sequencing held. The fault is found by the drain's *own* full blocking GC, in the | ||
| garbage of the mesh that had just been torn down. |
There was a problem hiding this comment.
Agreed, and fixed in 46e5f03. 'Garbage' is gone. The text now says the fault is in metadata built for the torn-down mesh, that a full blocking GC examines reachable objects too, and that reachability is unknown because the root walk failed.
| - Runs from 2026-10-07 17:40Z to 2026-10-08 ~06:00Z: **0** in 42. Of those, 28 used sets carrying | ||
| #6259 (`3.0.0-ci.10184`–`ci.10206`), all with Tiered PGO off, and each had FutuRe `exit 0`. | ||
| - At 1 %, 0.28 deaths were expected, so P(0) = 0.75. | ||
| - Separating a reduction from noise at 95 % takes **about 300** clean post-fix runs. |
There was a problem hiding this comment.
Agreed, and fixed in 46e5f03. The same correction as the other statistics thread: one-sided Fisher p ≈ 0.162 at 0/300, which does not establish a reduction. 300 runs is now defined as an operational threshold whose one-sided 95 % upper bound on the post-fix rate is ~1 % (0.994 %). The section also says a closure on that basis must state it, and that the issue reopens on the first death.
…ificant reduction; the pointer scan bounds only its window; no unreachability claim Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Docs only. This adds the readings of two
MeshWeaver.FutuRe.TestSIGSEGV dumps from 2026-10-07 toDoc/Architecture/DebuggingNativeCrashes, as sightings #19 and #20. The new section sits after #18. This change set conserves the work for Systemorph/MeshWeaver.Plugins#1605.What the section records
The two crashes.
3.0.0-ci.101253.0.0-ci.10157NT_SIGINFO11/1/0x011/1/0x010.0.12, build-id79945f51…)plan_phase+0x24fcmark_object_simple1+0x9c7Both read a MethodTable word that is zero.
What #19's zeroed object is. ClrMD names it as lock #1 of the STJ
PolymorphicTypeResolvercache forWorkspaceReference<object>, which belongs to a mesh'sJsonSerializerOptions.When it happened.
CollectibleUnloadDrain.WaitUntilCollectedAsync→GC.Collect(), afterDISPOSE_DONE.NodeAssemblyLoadContexts are Unloading.Why core #6259 and #6254 are not shown to explain it.
Rate before and after.
What stays open. The section ends with what is not established, and with the close criterion for Plugins#1605: about 300 clean post-fix runs, or a deterministic repro.
Verification
dotnet build -c Release -warnaserroronMeshWeaver.Documentation: 0 warnings, 0 errors.MeshWeaver.Documentation.Test: 0 warnings, 0 errors.MeshWeaver.Documentation.Test: 721/721 passed. This includesDocumentationLinkIntegrityTest; the new links are../NodeTypeCompilationand../CollectibleThreadStaticHandleReuse.No recycle is needed: this changes only an embedded doc page.
Pairs-with: none — docs only, no public surface removed
Implementers: none — no interface member added
Mirror-sync: none — no i18n catalog change
🤖 Generated with Claude Code