db/state: fix Close vs external MergeLoop/BuildFiles* WaitGroup races - #22203
Conversation
MergeLoop registers on the aggregator's WaitGroup from goroutines the aggregator did not spawn (e.g. the node's background-maintenance goroutine), so its Add was unordered against Close's Wait — WaitGroup reuse from zero, flagged by -race, and an unregistered merge loop could keep running through Close's file teardown. Guard the registration with a closing flag so every Add is either ordered before the Wait or refused. Same fix for ForkableAgg.
There was a problem hiding this comment.
Pull request overview
Fixes a sync.WaitGroup data race between Aggregator.Close()/ForkableAgg.Close() (wg.Wait) and externally-spawned MergeLoop() calls (wg.Add), by making MergeLoop lifecycle-aware and refusing to register once shutdown begins.
Changes:
- Add
closing+closingMutoAggregatorandForkableAgg;Closesetsclosingbefore waiting, andMergeLoopregisters under the same mutex (or returns early if closing). - Add race-regression tests that invoke concurrent
MergeLoopcalls across theClose()window for bothAggregatorandForkableAgg.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
| db/state/aggregator.go | Adds shutdown-aware guard around MergeLoop registration to prevent wg.Add racing Close’s wg.Wait. |
| db/state/forkable_agg.go | Mirrors the same shutdown-aware guard for ForkableAgg.MergeLoop vs ForkableAgg.Close. |
| db/state/aggregator_close_test.go | Adds regression test reproducing the Close-vs-concurrent-MergeLoop race scenario. |
| db/state/forkable_agg_test.go | Adds regression test reproducing the Close-vs-concurrent-MergeLoop race scenario for forkable aggregator. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Extract the closing flag into closingWaitGroup (TryAdd/BeginClose) shared by Aggregator and ForkableAgg; apply it to BuildFilesInBackground (both types) and BuildFiles2 alongside MergeLoop. BeginClose also replaces the unsynchronized ctxCancel idempotency latch, making concurrent Close safe. Red-first tests per door plus concurrent-Close; trim the MergeLoop race test (t.Parallel, 4 iterations) from 42s to 6-9s under -race.
Code reviewThe core fix is sound. The A few quality / test-coverage notes, ranked: 1. The ForkableAgg race twins drop the barrier + stagger their Aggregator counterparts rely on. 2. Nits:
|
- closingWaitGroup no longer embeds sync.WaitGroup: a named field plus TryAdd / AddFromRegistered / Done / Wait / BeginClose keeps a bare, latch-bypassing Add off the type's API. - Concurrent Close now blocks every caller until teardown completes: the caller that loses the BeginClose latch waits on a channel the winner closes via MarkClosed after closeDirtyFiles. - ForkableAgg Close-vs-* tests gain the start-barrier + per-goroutine stagger their Aggregator twins use, so the Add-during-Wait interleaving is actually produced. They stay serial (not t.Parallel) because setup registers into the global forkable Registry whose reads aren't lock-guarded, so parallel setups race there. - TestAggregatorCloseVsConcurrentBuildFiles2 exercises the merge-spawn path (doMerge=true) and drops a dead trailing wg.Wait.
Replace the hand-rolled closed-channel + BeginClose/MarkClosed protocol with a sync.Once-backed RunClose(teardown func()): Once.Do already runs teardown exactly once and blocks concurrent/later callers until it returns, which is the same block-until-torn-down, idempotent-Close guarantee — and matches the sibling Collector.Stop idiom called two lines into Aggregator.Close. Removes the forgettable "winner must MarkClosed" obligation; the closing latch that TryAdd checks is set inside the Once.
The function promises a channel closed when aggregation is done, but the doMerge==false branch returned without closing fin. Not a live bug — both callers pass doMerge=true — but the latent contract violation would deadlock a future no-merge caller waiting on the channel.
AddFromRegistered's contract was "the caller is an already-registered goroutine" — a precondition Go code can't observe, since any function may run inline or in a goroutine at the caller's choosing. Remove it and route every registration through the latched TryAdd, which depends only on the close latch, not on caller context: - The buildFiles / buildFile errgroup children dropped their wg registration entirely — g.Wait() already joins them before the (registered) caller returns, so Close covers them transitively; the extra count was redundant. - The fire-and-forget merge spawns now TryAdd: refused once closing, in which case the merge is skipped (and fin closed) instead of started during shutdown.
AskAlexSharov
left a comment
There was a problem hiding this comment.
i a bit modified to addresss Copilot's things.
@yperbasis feel free to review.
merging to fix main ci.
…nt data race on shutdown (#22244) **[SharovBot]** ## Problem A data race was detected in `TestImportClosesChaindataOnInitError` (race-tests CI job, 2026-07-03): ``` WARNING: DATA RACE Write at 0x00c0653acb90 by goroutine 321: github.com/erigontech/erigon/db/state.(*Aggregator).Close() db/state/aggregator.go:630 Previous read at 0x00c0653acb90 by goroutine 373: github.com/erigontech/erigon/db/state.(*Aggregator).MergeLoop() db/state/aggregator.go:1228 github.com/erigontech/erigon/node/eth.New.func16() node/eth/backend.go:1133 ``` ## Context PR #22203 (merged 2026-07-04) addressed this race by replacing `sync.WaitGroup` with a `closingWaitGroup` latch in `Aggregator`, making `MergeLoop`'s `TryAdd()` properly ordered against `Close()`'s `BeginClose()+Wait()`. ## This PR This PR provides an additional, complementary fix: track the MergeLoop goroutine in `bgComponentsEg` so `Stop()` → `bgComponentsEg.Wait()` explicitly waits for the MergeLoop goroutine to exit before `chainDB.Close()` is called. Without this, `bgComponentsEg.Wait()` in `Stop()` returns without waiting for the MergeLoop goroutine (since it was launched as a bare `go func()`), meaning the goroutine could theoretically still be running when `chainDB.Close()` begins. The `closingWaitGroup` handles the WaitGroup reuse race, but this PR makes the shutdown ordering explicit and unambiguous. **Changes:** - Moves the MergeLoop goroutine from a bare `go func()` to `backend.bgComponentsEg.Go()` - Fixes the typo: `"snapashot"` → `"snapshot"` in the error message - Filters context cancellation errors (expected on shutdown) from `logger.Error` ## Testing - `go test -race -count=10 ./cmd/utils/app/ -run TestImportClosesChaindataOnInitError` — all 10 runs PASS, no data race - `go build ./...` — succeeds - No test files modified Fixes CI: https://github.com/erigontech/erigon/actions/runs/28674562917/job/85045122311 Co-authored-by: SharovBot <sharovbot@erigon.tech> Co-authored-by: Giulio Rebuffo <giulio.rebuffo@gmail.com>
Fixes the data race that failed the
race-tests / tests-linux (other, serial)job on #22163's CI run (failing job):TestImportClosesChaindataOnInitErrorflaggedAggregator.Close'swg.WaitracingMergeLoop'swg.Add.Root cause
MergeLoop,BuildFilesInBackground(on bothAggregatorandForkableAgg) andBuildFiles2register on the aggregator's lifecycle WaitGroup from whatever goroutine calls them. For external callers nothing orders thatAddagainstClose'sWait. AnAddfrom a zero counter concurrent withWaitissync.WaitGroupreuse (undefined behavior, flagged by-race), and semantically the unregistered goroutine can keep running whileClosetears down the dirty files.The CI failure hit the
MergeLoopdoor (the background-maintenance goroutineeth.Newspawns), but the same race is reachable through the sibling doors, so this PR guards all of them:Aggregator.BuildFilesInBackground— live on a stock node: the fire-and-forget FCU background-prune goroutine (FcuBackgroundPrunedefaults to true) ends inCollateAndPrune → BuildFilesInBackground, and its adaptive budget can reach 2/3 slot ≈ 8s on mainnet whileEthereum.Stopwaits on the exec semaphore for at most 5s (WaitIdle) before proceeding tochainDB.Close → Aggregator.Close → wg.Wait. Same shape forProcessFrozenBlocks' commit cycle, which callsBuildFilesInBackgroundright after a ctx-oblivious MDBX commit.Aggregator.BuildFiles2andForkableAgg.BuildFilesInBackground— same unguarded caller-sideAdd; currently only reachable from tooling/tests, guarded all the same.#21528 fixed this class for the nested spawn sites by making the already-registered parent goroutine
Addbefore spawning; that pattern can't cover the entry points themselves — the registration has to be lifecycle-aware. The window became CI-visible when #22058 added a test that starts the import node, fails init, and immediately stops it.Fix
A small
closingWaitGroup(async.WaitGroupwith a mutex-guarded close latch), shared byAggregatorandForkableAgg:TryAddregisters unless the latch is set. Every external entry point (MergeLoop,BuildFilesInBackground,BuildFiles2) refuses once closing and returns its usual "nothing to do" result, like the existingdbg.NoMerge()/buildingFiles-CAS early-outs (the CAS is unwound andfinclosed on refusal).Closelatches viaBeginClosebefore cancelling the context and waiting. The mutex gives the happens-before edge: everyAddis either strictly ordered beforeClose'sWait, or refused.BeginClosedoubles as the Close-idempotency latch, replacing the previously unsynchronizedctxCancel == nilcheck /ctxCancel = nilwrite — two concurrentClosecalls used to race onctxCanceland could invoke a nil func;Closeis now safe to call concurrently with itself.TryAdd. The fire-and-forget merge spawns use it too (refused once closing: the merge is skipped andfinclosed, rather than an unconditionalAdd), and thebuildFileserrgroup children no longer touch the lifecyclewgat all (see Why dropping thewg.AddinbuildFilesis safe below).TDD
Red first, one test per door plus concurrent-Close, each reproducing the exact race signature under
-racebefore the fix:TestAggregatorCloseVsConcurrentBuildFilesInBackground—wg.Wait(aggregator.go:638) vswg.Add(aggregator.go:2171)TestAggregatorCloseVsConcurrentBuildFiles2—wg.Waitvswg.Add(aggregator.go:1167)TestAggregatorConcurrentClose— data races onctxCancel(reads ataggregator.go:627/633vs the nil-write at634)TestForkableAggCloseVsConcurrentBuildFilesInBackground—forkable_agg.go:462vs theBuildFilesInBackgroundAdd, plus the secondary merge-internals vscloseDirtyFilesraceTestForkableAggConcurrentClose—ctxCancelraces escalating to an actual nil-func-call panicGreen after the fix. Verified additionally:
TestAggregatorCloseVsConcurrentMergeLoop,TestForkableAggCloseVsConcurrentMergeLoop, both db/state: fix Aggregator & ForkableAgg Close vs background MergeLoop WaitGroup race #21528 regression tests, andTestAggregatorCloseReleasesBranchCache) pass under-racewith-count=2, zero race reportsTestAggregatorCloseVsConcurrentMergeLoopalso got cheaper: it burned ~42s under-race(~97% idle inWaitForFiles' 3-second poll ticker whenever a merge attempt overlappedClose); with 4 iterations andt.Parallelit runs in 6–9s and the whole Close suite finishes in ~19s under-race -count=2go test -short ./db/state/... ./db/kv/temporal/...passesmake lintclean (two consecutive runs),make erigon integrationbuildsTestImportClosesChaindataOnInitError(the original CI failure) passes locally without-race; it cannot run under-raceon darwin/arm64 at all (fatal error: too many address space collisions for -race modeat startup — a Go runtime limitation with this test's MDBX mappings, unrelated to this change), so the Linuxrace-testsjob on this PR is the definitive check for itWhy dropping the
wg.AddinbuildFilesis safebuildFiles(and forkablebuildFile) build their per-domain / per-II files on anerrgroup.Groupand block ong.Wait()before returning. Two facts make a separate lifecycle-wgregistration of those children redundant:buildFilesis only ever called synchronously from the background goroutine that already registered on the lifecyclewgviaTryAdd(thebuildFilesInBackground/BuildFiles2goroutine).defer wg.Done()untilbuildFilesreturns — i.e. untilg.Wait()has joined every child.So
Close'swg.Wait()already blocks on those children transitively, through the still-held count of the entry goroutine. Registering them on the lifecyclewgas well was pure double-counting — it changed the counter value but not the set of goroutinesClosewaits for. Removing it keepsClose's guarantee intact while getting rid of anAddwhose safety depended on the caller's context (which Go code can't observe — a function doesn't know whether it runs inline or in a goroutine).