execution/commitment: fix state reader leak across sequential batches - #22460
Merged
Merged
Conversation
AskAlexSharov
approved these changes
Jul 15, 2026
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Jul 15, 2026
AskAlexSharov
enabled auto-merge
July 15, 2026 12:01
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Jul 15, 2026
5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue
as seen from issue #22118, there's an intermittent assertion failure in the execution pipeline when running block execution in a loop over sequential batches (e.g. in
TestFromZero_BranchCacheCoherentAcrossBatches/parallel).The failure happens when initialising the next batch (Batch N+1). On startup,
ExecV3callsSeekCommitmentto restore the trie state. If we are in offline mode (inMemHistoryReads == false) at this stage, historical in-memory reads are not allowed. However, if the previous batch (Batch N) was running in parallel mode, the customasOfStateReaderof the commitment calculator (which usesGetAsOfreads) leaks on the shared commitment context.This stale state reader leak causes the restore path in the new batch to attempt historical GetAsOf reads, hitting the check in the memory overlay:
failed restore state: GetAsOf called on TemporalMemBatch with inMemHistoryReads offIt was timing dependent and flaky, because the leak only manifests when the shared context is reused across batch boundaries under specific goroutine scheduling conditions.
Fix
We fixed this by explicitly setting sdc.stateReader = nil in the
ClearRam()method of theSharedDomainsCommitmentContext.ClearRam()is the designated cleanup routine invoked at the end of each batch execution to flush the in-memory RAM overlays. By hooking the state reader cleanup into ClearRam() we ensure that once we are done with a batch, and close its transaction scope, any stale state reader reference is completely discarded.When the next batch starts and calls SeekCommitment, the context now has stateReader set to nil. This way, it can properly fall back to the regular LatestStateReader (which does normal, direct database reads) rather than trying to read via the time-travel overlay.
Tests
Since the existing suite already tests this behaviour, no new tests were added:
TestExec_RestoresCommitmentStateReaderchecks that the execution stage restores the commitment state reader cleanly.TestFromZero_BranchCacheCoherentAcrossBatches/paralleltests parallel mode sequential execution batches with a reusedSharedDomainsoverlay, verifying that the state of the commitment is cleanly restored across batch boundaries;Both tests now pass reliably when the race detector is enabled.
Closes #22118