Skip to content

[WIP] btinterp vs PrefixIndex on #21778 base — perf experiment - #21865

Closed
sudeepdino008 wants to merge 29 commits into
erigontech:mainfrom
sudeepdino008:exp/bt-compare
Closed

sudeepdino008 wants to merge 29 commits into
erigontech:mainfrom
sudeepdino008:exp/bt-compare

Conversation

@sudeepdino008

Copy link
Copy Markdown
Member

WIP / experiment — not for merge. Integrates three upstream branches to benchmark .bt navigation methods head-to-head on the new streaming .bt format.

Base: #21778 (alex/bt_streaming_build3_36) — streaming/footer .bt format, M-in-file, off-heap EF build.
On top:

Conflict resolution: kept #21778's footer format; ported PrefixIndex's Node.off as a runtime-only cache (not encoded); fixed the new BuildBtreeIndexWithDecompressor(existenceFilterPath) signature at stale call sites. Full btindex test suite passes (incl. TestInterpEquivBinary, prefix tests).

Bench (scratch/btcmp, M64, cold = vmtouch -e, 15k sampled keys, full Get)

SMALL — commitment 9085-9086 (1M, 0.5 GB):

method open cold µs cold faults/op warm µs heap
binary 2 ms 235 2.90 233 ~0 MB
btinterp 1 ms 141 1.75 140 ~1 MB
PrefixIndex 591 ms 256 3.21 258 ~4 MB

BIG — commitment 8192-8704 (141M, 40 GB), .bt = 235 MB (all three):

method open cold µs cold faults/op warm µs heap
binary 127 ms 175 2.06 168 87 MB
btinterp 136 ms 110 1.22 103 87 MB
PrefixIndex 52,434 ms 761 9.52 764 91 MB

Findings

  • Latency is ~entirely major faults (~80 us each); CPU/probe cost is in the noise.
  • btinterp best on cold (fewest faults via interp locality), trivial open.
  • PrefixIndex Open is O(.kv size) — a 2x full .kv scan (52 s on the 40 GB file). It warms the whole .kv, so if the entire .kv fits RAM lookups are ~0-fault/7 us (its 37-62% warm win). But with a genuinely cold .kv it is the worst (9.52 faults / 761 us — scatters more than binary), plus the catastrophic open.
  • .bt size and heap are comparable across all three.

Net: for big/cold commitment under memory pressure (the #21795 regime), btinterp is the clear winner; PrefixIndex only wins when the whole .kv stays resident.

awskii and others added 29 commits March 26, 2026 12:59
New PrefixIndex type — drop-in replacement for BpsTree lookups.
Per-prefix bucket architecture with adaptive node filling.

- [65536]prefixBucket{firstDI, endDI, nodes} — O(1) prefix lookup
- Adaptive: supplementary scan fills empty buckets with middle key
- narrowWithNodes: per-bucket binary search (max 8 nodes)
- Exact match shortcut: skip disk search on node cache hit
- L1 computed from L2 at build time

Benchmarks (same .kv file, random access):
  100K keys: Seek 53% faster, Get 62% faster
  1M keys:   Seek 46% faster, Get 37% faster

Co-Authored-By: Shuo <shuo@erigon.dev>
PrefixIndex is only built when ERIGON_USE_PREFIX_INDEX=true.
Default: false (BpsTree only, current behavior).

Usage: ERIGON_USE_PREFIX_INDEX=true ./erigon ...

Co-Authored-By: Shuo <shuo@erigon.dev>
Fix F3: Get returns (nil, false) instead of ErrBtIndexLookupBounds
for non-existent key in last bucket where endDI==count.

Fix F4: addNode copies key bytes via common.Copy to prevent
mutation risk from external nodes slice.

Add 11 correctness tests: boundary keys, exact match, non-existent,
concurrent reads, BpsTree comparison, node key stability.

Co-Authored-By: Shuo <shuo@erigon.dev>
…h/erigon into awskii/prefix-index-standalone
Lint fixes:
- Replace bytes.Repeat([]byte{0}, N) with make([]byte, N) (gocritic zeroByteRepeat)
- Replace string comparison with bytes.Equal (gocritic stringXbytes)
- Remove trailing newline (gofmt)

Bench fixes:
- Add missing buildBtreeIndex calls in BenchmarkPrefixIndexGet,
  BenchmarkSeekComparison, and BenchmarkGetComparison so the .bt
  index file exists before OpenBtreeIndexAndDataFile is called
- Add dbg.UsePrefixIndex=true in BenchmarkPrefixIndexSeek so
  bt.search (PrefixIndex) is initialized

Co-authored-by: shuo <shuo@erigon.dev>
…tandalone

# Conflicts:
#	db/datastruct/btindex/bpstree_bench_test.go
…tion

- Remove keyCmpFunc callback from PrefixIndex struct and constructors
- Add compareKey() method using seg.Reader.MatchCmp directly (zero-copy)
- Simplify Seek() binary search: no error return from compare path
- Fix comparison direction (cmp < 0 was inverted)
- Update tests to match new NewPrefixIndex/NewPrefixIndexWithNodes signatures
Search the leaf window by interpolation instead of binary search: estimate the
target from the bound keys' bytes after their common prefix. Probes stay
clustered near the target, so cold reads fault far fewer distinct .kv pages.
Default on (BT_INTERP=true); falls back to binary after BtInterpBudget (8)
probes. No index/format change; M unchanged.

Cold btnav us/op (M=256, mainnet 0-8192): storage 171->108, accounts 147->94,
code 258->109. The win is page locality, not fewer probes. Bound buffers are
stack-backed so the probe loop is allocation-free.

TestInterpEquivBinary (fixed keys) and TestInterpEquivBinaryVarLen (variable-
length, commitment-style) assert interp == binary for hits+misses across
budgets. Same change as erigontech#21794 (-> performance), here targeting main.
Add a footer-based .bt layout: body (nodes + EF) followed by a metadata
payload and a fixed 16-byte anchor (footer_len | flags | format_version |
magic), with magic last so opens fail fast on a wrong/truncated file.

Build the dense Elias-Fano offsets eagerly off-heap from KeyCount+MaxOffset
instead of buffering every key through ETL, and recompute each node's data
index as di=i*M rather than storing it. A non-zero leading byte distinguishes
footer files from the legacy EF-first format.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants