Skip to content

Flaky: TestFromZero_BranchCacheCoherentAcrossBatches/parallel — GetAsOf on TemporalMemBatch with inMemHistoryReads disabled #22118

Description

@mh0lt

Flaky test

TestFromZero_BranchCacheCoherentAcrossBatches/parallel — execution/execmodule/execmoduletester (added in #21952).

Run with the error: CI Gate run 28449207543, race-tests / tests-linux (ubuntu-latest, execution-other, serial). The job passed on retry within the same run (final job conclusion success), confirming it is intermittent.

Symptom

--- FAIL: TestFromZero_BranchCacheCoherentAcrossBatches (0.48s)
    --- FAIL: TestFromZero_BranchCacheCoherentAcrossBatches/parallel (0.24s)
        from0_genesis_internal_test.go:319:
            Received unexpected error:
            failed restore state : GetAsOf called on TemporalMemBatch with inMemHistoryReads disabled

Analysis

The error itself is a deterministic mode guard (TemporalMemBatch.GetAsOf returns it when inMemHistoryReads is false), so this is not value corruption — it's a timing race that surfaces as the guard when it loses:

  • The test re-executes from 0 in 1 KB batches via the integration path, which sets SetInMemHistoryReads(false) on the SD (cmd/integration/commands/stages.go).
  • Each batch calls SeekCommitment at the start of ExecV3 (exec3.go:122), with inMemHistoryReads still false. The restore path is SeekCommitment → restorePatriciaState → hext.SetState(...), which can issue a GetAsOf.
  • In parallel mode the previous batch's commit is asynchronous (commit goroutine + commitResults). If the prior batch's mem-batch flush hasn't completed when the next batch's SeekCommitment runs, the GetAsOf hits the still-populated TemporalMemBatch while inMemHistoryReads is false → the guard fires.
  • Parallel exec only enables inMemHistoryReads inside the executor (exec3_parallel.go:240, capturing/restoring the caller's value — added for parallel commitment calculations implemented #20805 against this same error string), which is after the start-of-batch SeekCommitment.

So the trigger is an ordering race between batch N's async commit/flush and batch N+1's pre-exec SeekCommitment restore. Timing-dependent → intermittent.

Local repro attempts (all green)

On mh/perf-statecache-lru-pr HEAD, execution/execmodule/execmoduletester:

  • -race -run TestFromZero_BranchCacheCoherentAcrossBatches: 6/6 pass
  • -race -count=12 (single test): pass
  • -race -count=10 at GOMAXPROCS=1, 2, 4: pass
  • full package -race, 4×: pass

Could not reproduce locally; it appears to need the concurrent-package CPU pressure of the race-tests group to widen the window.

Notes

cc parallel-exec / commitment owners.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions