Skip to content

blocks files: move refcnt from per-file to per-visibleFiles object (like state files) - #22246

Merged
awskii merged 10 commits into
mainfrom
alex/blk_rc_36
Jul 9, 2026
Merged

awskii merged 10 commits into
mainfrom
alex/blk_rc_36

Conversation

@AskAlexSharov

@AskAlexSharov AskAlexSharov commented Jul 5, 2026 •

Copy link
Copy Markdown
Collaborator

porting the Aggregator's visibleFiles's refcount (lock-free, generation-chained bundle-refcount) file
reclamation to the block files (db/snapshotsync)

What

Replaces the per-file two-atomic reclamation on the shared DirtySegment
(refcount atomic.Int32 + canDelete atomic.Bool, with RoTx.Close doing
if refcount==0 && canDelete { closeAndRemoveFiles() }) with an MVCC-style,
oldest-reader-watermark scheme, exactly as db/state's Aggregator already does:

  • snapshotVisible becomes a generation node (refcnt, retired, next), forming an
    oldest→newest chain; RoSnapshots.oldestVisible is the reclaim head.
  • Readers pin the whole bundle once (validate-after-pin in acquireVisible), instead of
    per-file increments — closes the visible.Load()→pin window.
  • Physical deletion is gated on the watermark: a retired file is unlinked only once every
    generation that referenced it has drained. No file is unlinked while a reader still
    mmaps it (fixes the eager-unlink hazard in RemoveOverlaps, incl. the Windows
    mmap-unlink case).

Design notes

  • Single lock, matching the Aggregator's one dirtyFilesLock: dirtyLock guards both
    dirty and the generation chain. recalcVisibleFiles(alignMin, retired) is called with
    dirtyLock held, so a mutator's dirty change and the bundle publish are one atomic step.
  • RemoveOverlaps mirrors cleanAfterMerge: opens a View, retires the subsumed
    segments under the lock, and the real unlink (each segment's .seg + indexes) happens
    off-lock at View.Close — or defers to the true watermark if another reader still holds
    the generation. No manual removeOldFiles of tracked files.
  • Caplin stores (CaplinSnapshots, CaplinStateSnapshots) construct every segment
    frozen, so the per-file machinery was already inert there; they only migrate off the
    removed atomics (a dead canDelete guard drop). No generation chain added to the
    append-only frozen stores.

Tests

  • snapshots_race_test.go: deferred-unlink-while-View-open, pending-retired protection,
    and a -race readers-vs-retire stress that asserts the chain collapses and no files leak
    after drain.

db/snapshotsync/..., freezeblocks, polygon/heimdall, polygon/bridge pass under -race.

@AskAlexSharov AskAlexSharov changed the title [3.7] [wip] db/snapshotsync: lock-free bundle-refcount snapshot file reclamation [3.7] [wip] blocks files: move refcnt from per-file to per-visibleFiles object (like state files) Jul 7, 2026
@AskAlexSharov AskAlexSharov changed the title [3.7] [wip] blocks files: move refcnt from per-file to per-visibleFiles object (like state files) [3.7] blocks files: move refcnt from per-file to per-visibleFiles object (like state files) Jul 7, 2026
db/snapshotsync conflict resolution: main's #22172 extracted the per-segment
`DirtySegment.refcount` close model into shared helpers and wired caplin to
them. This branch removed per-segment refcount in favour of the generation
model (snapshotVisible.refcnt), so the RoSnapshots close path keeps
detachNotInList + generation reclamation, while the shared
CloseSegmentsNotInList/closeAndDropNotProtected helpers are retained (they
dedup caplin's own close code) minus the refcount check, matching caplin's
existing no-per-reader-protection behaviour. Kept main's generic
FindOpenSegment/ClassifyOpenErr helpers and the OpenList dedup.
Close() deferred recalcVisibleFiles, so the empty generation was published
last — after the refcnt check and the segment-close loop. Since View() pins
the current generation lock-free (acquireVisible's hazard-pointer re-check),
a concurrent View() could load the still-current outgoing generation,
increment its refcnt, pass the re-check (s.visible unchanged because the swap
was deferred), and be handed segments Close was closing → use-after-close.

Publish the empty generation synchronously before reading the outgoing
generation's refcnt (captured as prev): a concurrent View() that pinned the
outgoing generation now fails its re-check and retries onto the empty one,
and any reader that pinned it earlier is counted in prev.refcnt, which we
honor. Matches the synchronous mutate-then-publish order Delete already uses.
@AskAlexSharov AskAlexSharov changed the title [3.7] blocks files: move refcnt from per-file to per-visibleFiles object (like state files) blocks files: move refcnt from per-file to per-visibleFiles object (like state files) Jul 8, 2026
@awskii
awskii requested a review from Copilot July 9, 2026 02:31

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@awskii
awskii added this pull request to the merge queue Jul 9, 2026
Merged via the queue into main with commit 85d5606 Jul 9, 2026
176 of 178 checks passed
@awskii
awskii deleted the alex/blk_rc_36 branch July 9, 2026 05:26
AskAlexSharov added a commit that referenced this pull request Jul 9, 2026
…e_36

Picks up blk_rc_36 now merged to main (#22246) plus Gloas CL (#22091). The three
db/snapshotsync conflicts were purely the #22343 rename (this branch renamed
snapshotsync.RoSnapshots -> BaseRoSnapshots; main's finalized blk_rc_36 kept the
old name): took main's canonical reclamation code (which already includes the
Close TOCTOU fix) and re-applied the rename in snapshots.go, merger.go and
snapshots_race_test.go.
awskii added a commit that referenced this pull request Jul 22, 2026
Main #22246 landed its own EL refcounted visible-generation retirement
core (snapshotVisible), so adopt it for the block path as-is and drop this
branch's generic visibleGenerations[P] and int CaplinStateType.

Caplin keeps the reader-safe retirement it added — RemoveOverlaps defers
the unlink while a live view pins the generation, and Close is drain-gated
so it never closes fds an older pinned view still references — but now as a
concrete caplin-local generation core over main's string-keyed model,
preserving #21901's per-type dump planner and the string type API.
awskii added a commit that referenced this pull request Aug 6, 2026
…ments input

Close read the refcount of the immediately-outgoing generation only, but
generations share *DirtySegment values and reclaimRetiredLocked walks the whole
oldest->current chain precisely because an older one can still be pinned. A
reader holding G1 across an intervening republish of G2 therefore had its
segments closed under it: Decompressor and indexes went nil, so reads returned
nothing or panicked in MakeGetter. Gate the close on the whole chain instead.

The defect predates this branch (the gate came in with #22246) and hits EL
blocks and bor/heimdall the same way; caplin riding BaseRoSnapshots just widens
the exposure.

openSegments scheduled one OpenIdxIfNeed per name, so a repeated name raced two
goroutines on the same segment's index slice - a torn slice header plus a
double recsplit.OpenIndex whose loser leaks. Deduplicate at that single funnel,
which covers OpenList, OpenFolder and OpenSegments alike. No current caller
passes duplicates.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants