exec: stop re-hashing bytecode on every state-object load - #22956
Conversation
getStateObject and reconstructCellFlags recomputed keccak(code) to keep the rebuilt stateObject's CodeHash in step with the code cells. The hash was already known: the version map and the write set store accounts.Code, which carries it, and a value read back from committed storage is covered by the account record's own CodeHash. Under noMaterialize (#22409) the state object is rebuilt on every getStateObject call, so a per-load hash became a per-read hash of the whole contract. In a 300s mainnet parallel-exec profile this was 15.01s -- 83.6% of all keccak in the process, 5.5x the KECCAK256 opcode itself (2.73s) -- reached from balance, nonce and code-hash reads alike. refreshCode now returns the writer's hash alongside the bytes. It cannot return accounts.Code: that type documents Hash == Keccak256(Bytes), and the read-set source stores bytes only. The new refreshedCode carries no such invariant and resolves the hash per source. BenchmarkGetStateObjectAfterCodeRead (ns/op): code size before after 32 B 331 175 -47% 1024 B 1_255 171 -86% 24576 B 23_620 171 -99.3% Before scales with code length; after is flat.
There was a problem hiding this comment.
Pull request overview
This PR optimizes IntraBlockState state-object reconstruction by avoiding repeated keccak256(code) hashing when code bytes are already available together with a trustworthy hash (or when the committed account’s CodeHash is authoritative). This fits into Erigon’s execution/state fast-paths for version-map backed (parallel/BAL) execution and noMaterialize operation.
Changes:
- Extend CodePath reads to carry an optional “known hash” alongside code bytes, and add
refreshedCode+ hash-resolution logic to reuse it. - Update
getStateObjectandreconstructCellFlagsto use the resolved hash instead of always hashing bytecode. - Add a benchmark and a regression-style test covering the state-object rebuild path after a prior code read.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
execution/state/read_paths.go |
Introduces refreshedCode and resolves code hashes without re-hashing when a known hash or authoritative committed CodeHash is available. |
execution/state/intra_block_state.go |
Switches state-object reconstruction to use refreshedCode.codeHash(...) rather than unconditional bytecode hashing. |
execution/state/code_rehash_bench_test.go |
Adds a benchmark and a test around noMaterialize state-object rebuild behavior after warming CodePath. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
The previous test seeded an account whose CodeHash already equalled keccak(bytes), so it passed with or without the change. Replaced with a subtest that seeds a CodeDomain entry disagreeing with the account record: the account hash must win. It fails on main (the rebuilt object takes keccak(stale bytes)) and passes here. Also assign mapCodeKnownHash wherever mapCodeVal is assigned, so the two cannot drift, and switch the benchmark to b.Loop().
…mment Both discriminating tests now exist: the committed path (account record wins over stale CodeDomain bytes) and the version-map path (the cell's own hash wins over keccak of its bytes). The second one guards the mapCodeKnownHash wiring -- a mis-plumbed field there would hand back a wrong hash, not just a slow one. Both fail on main. The comment above the resync still described re-deriving the hash from the bytes, which is no longer what the code does.
|
before: #22955 (comment) |
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 3 out of 3 changed files in this pull request and generated no new comments.
Suppressed comments (1)
execution/state/code_read_parallel_test.go:127
- Benchmark timing currently includes the one-time setup cost (committedCodeIBS + warm GetCode), which can skew ns/op and makes results less comparable to other benchmarks in this repo that call b.ResetTimer() before the measured loop.
b.Run(fmt.Sprintf("code=%dB", codeLen), func(b *testing.B) {
ibs, addr := committedCodeIBS(b, codeLen, nil)
b.ReportAllocs()
for b.Loop() {
if _, err := ibs.getStateObject(addr, false); err != nil {
b.Fatal(err)
}
}
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 3 out of 3 changed files in this pull request and generated no new comments.
Suppressed comments (2)
execution/state/read_paths.go:1116
- The comment says "KnownHash is Nil", but this field is an accounts.CodeHash, so the sentinel value is NilCodeHash. Using the constant name here avoids confusion with Go's nil.
// refreshedCode is refreshCode's result. Unlike accounts.Code it promises no
// Hash == Keccak256(Bytes): KnownHash is Nil when the source stored bytes only.
type refreshedCode struct {
Bytes []byte
KnownHash accounts.CodeHash
}
execution/state/code_read_parallel_test.go:128
- The benchmark includes setup work (committedCodeIBS + warming GetCode) in the timed section, which can skew ns/op—especially for small code sizes. Most benchmarks in this package reset the timer after setup (e.g. execution/state/reset_bench_test.go:45).
b.Run(fmt.Sprintf("code=%dB", codeLen), func(b *testing.B) {
ibs, addr := committedCodeIBS(b, codeLen, nil)
b.ReportAllocs()
for b.Loop() {
if _, err := ibs.getStateObject(addr, false); err != nil {
b.Fatal(err)
}
}
})
yperbasis
left a comment
There was a problem hiding this comment.
Verified locally: both new tests fail on main and pass here; the benchmark reproduces (flat ~225ns vs scaling to ~27us at 24KB); execution/state suite green.
Two non-blocking nits:
-
The residual-keccak list in the description misses one case: for a dirty address (
journal.dirties) the read-once gate is bypassed, so a committed code read resolves through theMVReadResultNonematching-header branch with sourceReadSetReadand still re-hashes on every rebuild. Unimproved rather than regressed — maybe worth a sentence in the description. -
In the
MVReadResultDoneread-set hit (prHeader.Version == hdr.VersioninversionedReadCore), the just-probed cell hash is already inr.hashOfMapCodeVal, butrefreshCodediscards it and the caller re-hashes. Returning it foroutcomeReadSetHitis safe for every source of that outcome (the other read-set-hit branches leave the field Nil, and version equality pins the same CodePath cell), so it would shave part of the remaining keccak without the CodePath/CodeHashPath pairing assumption of the deferred suggestion. Follow-up material.
A dirty address bypasses the read-once gate, so the rebuild re-probes the version map and resolves through outcomeReadSetHit. The probed cell's hash was discarded there and the caller re-derived it from the bytes.

getStateObjectandreconstructCellFlagsrecomputedkeccak(code)on every rebuild, to keep the rebuilt stateObject'sCodeHashin step with the code cells. The hash was already available at every source that produced the bytes.Why it is hot
Erigon's CodeDomain is keyed by address, not by code hash (geth/reth/gevm all key their code store by hash, so they never re-derive it on a load). Two independent cells — the account record's
CodeHashand the code bytes — can therefore disagree, e.g. when a prior tx in the block sets an EIP-7702 delegation whoseAddressPathrecord was published with the old hash. The recompute was a blunt way to makeobj.codeandobj.data.CodeHashagree.That was affordable while
getStateObjectcached its result. UndernoMaterialize(#22409) it no longer does — the object is rebuilt on every call — so a per-load hash became a per-read hash of the entire contract, reached from balance, nonce and code-hash reads alike.In a 300s mainnet parallel-exec profile:
getStateObject->Keccak256HashopKeccak256(theKECCAK256opcode)The internal resync cost 5.5x the opcode it exists to implement.
Fix
refreshCodenow returns the writer's hash alongside the bytes, and the caller resolves it per source:accounts.Code, which already carries the hash.CodeHashis authoritative; the bytes cannot be newer than it. This produces exactly the pairingstateObject.Code()builds on its lazy path (accounts.Code{Hash: so.data.CodeHash, Bytes: code}).One case still hashes on every rebuild: a dirty address reading committed code.
journal.dirtiesbypasses the read-once gate, so the rebuild re-probes, misses the version map, and resolves through the matching-header branch with sourceReadSetRead— which does not carry the account record's authority the wayStorageReaddoes. Unimproved rather than regressed; resolving it means widening the account-record rule to a read-set source, which is a behaviour change of its own and belongs in a separate PR.refreshCodecannot just returnaccounts.Code: that type documentsINVARIANT: Hash == Keccak256(Bytes), and the read-set source stores bytes only. The newrefreshedCodecarries no such promise, so a caller cannot silently take a bogus hash for a real one.Numbers
BenchmarkGetStateObjectAfterCodeRead— the state-object rebuild every account-field read falls through to. n0 (EPYC 4344P), interleaved binaries, benchstat n=8, all p=0.000:Before scales with bytecode length; after is flat. Allocations identical — this removes compute, not garbage.
Behaviour
One deliberate change: when the CodeDomain entry disagrees with the account record, the account record now wins. The CodeDomain is keyed by address, so it can hold bytes the account no longer owns — a cleared 7702 delegation leaves them behind — and the old code let those bytes overwrite
obj.data.CodeHash/obj.original.CodeHash.stateObject.Code()'s lazy path already preferred the account record.TestCommittedCodeHashComesFromAccountRecord,TestPriorTxCodeWriteHashComesFromTheCellandTestPriorTxCodeWriteHashSurvivesReadSetHitall fail on main. Each gives its cell a hash that disagrees withkeccak(bytes), which is what makes the source of the hash observable — a real cell never lies, so on real input this changes cost, not values.