You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This epic owns the one-ingest multicore outcome and the direct throughput floor. #1448 owns the bounded core-scaling work and evidence; this epic is the canonical close gate for both outcomes. #1478's separate planning issue duplicated this plan and is superseded by this section; its revisions 10–14 remain archived at docs/development/evidence/ingest-floor-plan-1478-rev14.md. GitHub's issue relationships carry execution state and prerequisites. #735 is the sole M5 closing tracker.
Contract and decisions
The rate floor is 1,000,000 complete ingested edges per second at every supported scale: T_complete_ingest(E) <= E / 1,000,000. The boundary includes registration through publication and acknowledgement. The separate multicore outcome is a material speedup for one complete ingest as more cores are made available, measured under #1448's predeclared core-use criterion on a supported durable-project host. S26 admission, a faster sub-stage, or a single fast machine does not substitute for either outcome. No macOS run is required to measure the core-count speedup; do not claim an unmeasured Apple-specific rate.
The direct end-to-end floor and perf(storage): make one complete ingest scale with available cores #1448's one-ingest core-count scaling criterion are distinct acceptance outcomes. A fast single-core run cannot prove multicore use; a positive scaling curve cannot prove the 1M floor. Per-resource figures remain policy allocations and diagnostics, judged at matched boundaries; no CPU-per-edge ceiling is inferred from N-process probes.
Remove bytes and work ahead of and alongside structural experiments; try overlap before decomposition. Do not build a pipeline to hide preparation work.
The retained S18–S22 baseline is b6ffb088 from the #1530 archive. It measured S22 at 126,099 edges/s, a 7.93× gap to the floor; rung receipts and the manifest are linked above in the retained-evidence record. A later #1481 integrated-tree checkpoint (2026-09-22, aaf20d83 era) measured complete-ingest medians of 31.818 s at S18 (7.586 µs/edge; 7.59× allowance) and 547.803 s at S22 (8.163 µs/edge; 8.16× allowance). Its publication-repair A/B did not demonstrate a throughput gain. This S18/S22 checkpoint supplements, but does not replace, the archived S18–S22 ladder or the exact frozen SHA required for #1448.
perf(storage): make one complete ingest scale with available cores #1448 owns useful core-count scaling for one complete ingest. Partition-level shaping is its first bounded hypothesis, with the predeclared 10% whole-ingest S18/S20 adoption gate and unchanged digests/contracts. Its final evidence includes a matched core-count speedup curve and the limiting region; shaping alone cannot close the 1M/s gap.
The older detailed acceptance checklist below is retained for history. Under D1, historical fixed CPU-per-edge, serialized-fraction, core-count and byte ceilings are diagnostics rather than acceptance thresholds. #1448's predeclared complete-ingest core-count scaling criterion is the current multicore acceptance outcome. Close this epic on both that criterion and the rate at every supported rung, plus the no-degradation shape criterion, correctness/recovery/determinism, retained evidence, and any ADR required by a changed storage or authentication design. Keep #1476 regression ratchets active.
Any need for a process-wide admission manager, manifest format, or publication-protocol change stops the experiment that found it and routes to a separate decision under #1509. Preserve recorded UUID splitters, filesystem authority, total-memory admission, cooperative cancellation, durability, and the complete-ingest boundary. The full design cautions and reversibility notes from #1478 remain in its revision history and archived rev14 document.
#1572 records the maintainer decision to defer the unmet 1,000,000 complete edges/s floor to post-v0.6.0. This issue remains open with its rate target and separate multicore-scaling acceptance outcome unchanged; neither is marked passed or waived. The measured S22 result was 184,749 edges/s, and the current evidence does not support expecting recent changes to close the 5.41× gap. The later eight-core S18 curve measured 1.39× one-core throughput and 1.19 effective cores, so multicore scaling is also retained as an explicit unresolved post-release outcome.
#1387 is no longer an M5 child/blocker and is no longer assigned to the M5 milestone. This issue remains the post-release owner; any future performance experiment requires a bounded issue and coordinated quiet-host allocation. v0.6.0 still requires #900/#745 S26 lifecycle certification and all other approved release gates. #1476 regression ratchets remain active.
Retained evidence (#1530 / PR #1534). The archive index distinguishes complete rung evidence from surviving controller summaries. The b6ffb088 baseline retains S18–S22 rung/result JSON, receipts, plan, projection and controller summary; its MANIFEST.sha256 digest is 807609f97477e91c3cb5321261a4e9274f51573c150da88b320baf3b55c8342d. These links pin the merged archive commit from #1534. Historical raw evidence for 93c041df, f80f69fe (including S24), 1955f17d, ab1a713e, 4fbfe84c and 403fc02a was lost on 2026-09-21. Retained controller summaries support only their reported values; other historical numbers rest on the copied issue tables and comments, not missing per-operation receipts or publication-region data. The archived rev14 §12 is historical; this retained index is the current evidence location.
Superseded framing (2026-09-21, #1478 rev 13 D1 / rev 14). Statements below that the CPU budget is "banked" and the remaining distance is a "serial fraction" (≤2% or ~10% on 16 threads) are withdrawn: effective cores below one is a net CPU-time deficit, not a measured serial fraction, and no CPU-per-edge necessity is established. The floor, the direct deadline T_ingest(E) <= E / 1,000,000, and the acceptance criteria stand. Latest ladder results and the 2026-09-21 evidence loss are in the comments; retention rule #1530.
⚠️ Status, 2026-09-18 — read before the body below
Current plan and closing tracker (2026-09-19, rev 13): #1478.#1387 owns the 1M edges/s acceptance requirement; #1456 remains its construction-foundation child. #1478 coordinates their closure; neither closes automatically or substitutes for the other's acceptance evidence. M5's canonical tracker remains #735.
Current decisions supersede the historical diagnosis below. D3 keeps #1465 on encode for seam viability; a pass permits a bounded shaping experiment, not migration. Remove the ≥2× parallelism condition on both surfaces. The N-process 3.54× throughput result is not effective cores or a single-ingest ceiling: CPU ≤3.54 µs/edge and the mandatory 2.30× cut are withdrawn. #1476 preserves valid regression gates and the direct floor without that derivation. #1477 distinguishes CPU concurrency, scheduler states and speedup; #1480 attributes synchronization/scheduling/CPU effects and does not treat loop devices on one array as storage isolation.
Publication is owned by #1481, a native child and blocker of #1387. Its historical ~0.42 µs/edge scales with input and must be re-measured after CSR publication landed; budget it inside complete ingest, not as fixed startup. D2's authorized ordering is recorded in #1478; no new performance measurement is claimed by this update.
Historical material follows. Claims below that the CPU budget is banked, the surrogate chain is the critical bottleneck, or foundation-first remains the current order are superseded by #1478 rev 13. Preserve the recorded measurements and original scope, not those withdrawn conclusions.
The floor is binding. Confirmed by the maintainer; no invariant is exempt from being re-examined against it. S26 admission does not close this epic.
Three numbers in the body below are superseded, all by same-rung measurements on f80f69fe (clean-f80f69fe-evidence/):
body says
measured now
72,989 edges/s at the reference scale (S22, fa6447cc)
104,390 edges/s, +43.0%
"a fourteen-fold improvement"
9.58x
11.93 µs CPU/edge, budget "under 9.0"
8.46 µs(est.) — budget already met
Consequence: the joint constraint is now a single constraint. The CPU half is banked, so the remaining distance is almost entirely serial fraction. The body's "serialized fraction at most 2%" was written against the old baseline; with the CPU cut banked the floor needs roughly 10% on 16 threads (~3.7% on 8 physical cores). Today it is ~68–80%.
The serial region is a data dependency, not an I/O stage:validate_staged_details → assign_surrogates → resolve_endpoint_surrogates, the 48–50% remainder that capped #1448. Filed as #1464.
Sequencing: foundation before optimisation — #1462 (instrument) → #1463 (resource policy) → #1439 (skew) → #1465 (seam spike) → #1464 (prefix sum) → #1448 (seal) → byte amplification under #1194.
The rss_bounded_or_plateaued gate was removed from admission on 2026-09-18; rss_growth_fraction is still reported. S25 is now gated on io_reader_publication_headroom alone.
Full detail in the comments below. The body is kept for its reasoning and its comparable-systems table, both of which still stand.
Goal
Make GraphForge ingest fast enough that importing a graph stops being something a user plans their day around.
Floor: 1,000,000 edges per second, sustained at every supported scale.
This is a floor, not a target. Two consequences follow, and both are stronger than a target would imply.
It must hold at every scale, not at one reference point.Our current defect is that throughput falls as the graph grows, from about 79,000 edges per second at 8 million edges to 73,000 at 67 million.Updated 2026-09-18 (f80f69fe): the shape is now flat through S22 — 101,175 / 105,415 / 105,802 / 104,390 edges/s at S18–S22 — but drops 6.6% at S24 to 97,469, with CPU/edge rising 8.46 → 9.12 µs. The drift is real and has moved to the top of the ladder. A floor forbids that shape outright. Meeting the number at one convenient size and drifting below it at a larger one is a failure, not a partial success.
The design point must sit above it. You do not engineer to your floor; you engineer above it so that variance, larger inputs and future changes do not breach it. Currently 72,989 edges per second at the reference scale, so the floor is a fourteen-fold improvement — superseded 2026-09-18. Same rung, same metric on f80f69fe: 104,390 edges/s, +43.0%. The floor is a 9.58x improvement, and the design point must be more.
What guaranteeing the floor requires
Ingest spends 11.93 microseconds of CPU per edge at 67 million edges, measured on fa6447cc — now 8.46 µs(estimated) on f80f69fe, a 29% cut. It still runs on less than one of 16 cores: 0.85–0.89 effective cores, flat across a 64x edge range. CPU consumed per edge is independent of contention, so it is the stable figure to reason from; wall-clock times in this run were inflated by other work on the host. Throughput is governed by how much of the path stays serial and how much per-edge work is removed. Amdahl's law on 16 cores:
Serial fraction
CPU cut
Design point
Headroom over floor
0%
25%
1,557,097
1.56x
1%
25%
1,353,997
1.35x
2%
25%
1,197,767
1.20x
2%
35%
1,382,039
1.38x
2%
50%
1,796,650
1.80x
5%
35%
1,026,657
1.03x, thin
5%
50%
1,334,655
1.33x
Read that table as a joint constraint rather than a menu. A 5% serial fraction reaches the floor only barely, and only with a substantial efficiency gain. A 2% serial fraction clears it comfortably. Comfortable headroom across the board needs the serialized region at or under 2% and per-edge work cut by at least a quarter.
This is more achievable than the first version of this table showed, because the baseline improved. Two storage changes landed after the evidence this epic was originally written against, and they cut CPU per edge from 15.14 to 11.93 microseconds, a 21% reduction. The floor did not move; the distance to it shortened.
That turns two vague instructions into measurable budgets:
Serialized fraction of the ingest path: at most 2%, targeting under 1% → at most ~10% on 16 threads (~3.7% on 8 physical cores).Relaxed 2026-09-18 because the CPU budget below is already met; the 2% figure was derived against the old baseline. Today the serial region is ~68–80%. This is the same rule the surveyed systems converge on, that the serialized region should contain only the metadata commit. No instrument computes it — see observe(storage): nothing measures the serial fraction of ingest, and #1387 budgets against it #1462.
Databend is the sharpest comparison: same core count as our bench host, running at roughly twice our floor. Setting the floor below it is deliberate, since a graph edge carries identifier mapping and adjacency construction that an analytical row does not. The first draft of this epic set 400,000, anchored to the Neo4j figure, which is both a weaker peer and a result from 2016 hardware.
Why this is a product problem, not a benchmark problem
Translated into what a user waits for, at our best measured rate:
Graph size
Today
At target
10 million edges
2.7 minutes
10 seconds
100 million edges
27 minutes
1.7 minutes
1 billion edges
4.9 hours
18 minutes
Two things make the lived experience worse than those numbers suggest.
Throughput degrades as the graph grows. Measured across a sixteen-fold range, edges per second fall about 24%. Users do not read that as a scaling curve. They read it as a tool that gets slower the more they put in it, which is the opposite of the impression a graph engine should give.
The process uses 0.85–0.89 of 16 cores (0.92 when this was written). A user watches one core pinned and fifteen idle. That reads as a broken tool regardless of the throughput number.
What this epic does and does not deliver
Ingest is 67% of lifecycle wall time, so even an infinite ingest speedup caps the total lifecycle improvement at about three-fold. That ceiling should be stated plainly, because "blazing fast ingestion" invites a broader reading than this epic delivers.
At a billion edges, with ingest at the floor and no other phase changed:
Time
Ingest at the floor
18 minutes
Every other phase
2.42 hours
Full lifecycle
2.71 hours
Down from roughly 7.8 hours today, a 2.7-fold lifecycle improvement from a sixteen-fold ingest improvement.
The next bottlenecks are already visible and are out of scope here: reopen proof at 0.98 hours and query at 0.72 hours, projected to a billion edges. Together they are 1.7 hours and are untouched by this work. They deserve their own epic once ingest lands.
Two caveats on the 2.71 hours. It assumes the floor holds at a billion edges, which is sixteen times beyond the reference scale and has never been measured. And it assumes the other phases stay linear, which they have been through 67 million edges but which nothing guarantees beyond it.
Baseline
Ingest is 58 to 67% of total lifecycle wall time. Measured on fa6447cc, the ladder run completed 2026-09-17:
Rung
Edges
Ingest
Edges/sec
S18
4,194,304
54.7 s
76,663
S19
8,388,608
105.6 s
79,431
S20
16,777,216
214.8 s
78,110
S22
67,108,864
919.4 s
72,989
Superseded 2026-09-18 by clean-f80f69fe-evidence/. Ingest wall and edges are measured; effective cores is rung-wide CPU/wall and µs/edge is derived from it, so that column is estimated:
Rung
Edges
Ingest
Edges/sec
Eff. cores
µs CPU/edge
S18
4,194,304
41.5 s
101,175
0.85
8.43
S19
8,388,608
79.6 s
105,415
0.86
8.15
S20
16,777,216
158.6 s
105,802
0.86
8.13
S22
67,108,864
642.9 s
104,390
0.88
8.46
S24
268,435,456
2754.0 s
97,469
0.89
9.12
At the reference scale: 265 bytes retained per edge against 4,579 bytes read for authentication, a constant factor of 17.3 that does not grow with size. Physical reads 858.3 GB, physical writes 363.8 GB, peak resident memory 196 MB.
What changed since the previous evidence
Two storage improvements landed between the earlier ladder run and this one. They are real and they are measurable at scale, though they leave the structure untouched.
S22 metric
Before
After
Change
Ingest
1,104.1 s
919.4 s
−16.7%
Ingest throughput
60,781 edges/s
72,989 edges/s
+20.1%
CPU per edge
15.14 µs
11.93 µs
−21.2%
Peak resident memory
262 MB
196 MB
−25.3%
Physical read
885.8 GB
858.3 GB
−3.1%
Retained storage
17.81 GB
17.81 GB
unchanged
Transient peak
36.76 GB
36.76 GB
unchanged
The smaller rungs show no difference, so these changes target work that only appears at larger merge fan-in. Reads fell 3%, so the roughly 17-fold read amplification stands, and retained and transient storage are identical to the byte. The diagnosis below is unaffected; only the starting point moved.
Wall-clock figures for the non-ingest phases in this run ran 11% to 23% above the previous one, and whole-run core utilisation fell from 0.92 to 0.87, because other work shared the host. Those timings should not be compared across runs. Counters and per-edge figures are unaffected by contention.
Diagnosis
External research across eight comparable systems built on Arrow, Parquet and DataFusion produced a clear picture of where we differ.
We read data back in order to hash it. Nobody else does. InfluxDB 3, GreptimeDB, Iceberg, Delta, Databend, Lance, ParadeDB and SlateDB were all examined. Not one performs a read-back verification pass over data it just wrote. Where integrity exists at all it is a non-cryptographic checksum computed inline by the pass that produces the bytes, verified lazily at read time or not at all. Several hash nothing on the data path.
The serialized region contains work that does not belong there. All eight converge on the same architecture: many parallel encoders funnelling into one cheap serialized commit. The design rule that follows is that the serialized region should contain only the metadata commit, never encoding, sorting, hashing or copying. GreptimeDB measured an 81% gain from enforcing exactly that boundary. DuckDB and DataFusion each measured roughly six-fold from parallelising their write paths.
The hash primitive is not the lever. This host has the SHA-256 hardware extensions and the profile implies we already hash at about 2.65 GB/s. Switching primitives buys roughly 6% of ingest. The number of passes is the problem, not their cost per byte.
Our merge is already good. Merge write amplification is about 1.4x, better than any published log-structured design. The 21x total write amplification is therefore not the merge, and compaction redesign is not the answer.
Workstreams
Ordered by expected impact. Each becomes a child issue.
Eliminate the redundant authentication passes.perf(storage): eliminate the redundant authentication passes that dominate ingest #1384, already filed. Hash once as bytes stream to their destination. Target 310 GB of hashed bytes down to 25 GB. Move any "is the store still intact" check behind an explicit administrative verify command rather than running it during ingest.
Attribute and remove the unexplained write volume. The merge writes 25 GB; the rung writes 372 GB. Staged chunks and the content-addressed install copy are the candidates. For calibration, a full log-structured store on object storage measures 3.0x with segment sealing and 5.5x without; we are at 21x.
Determine whether a global ordering requirement is serialising us. DuckDB's 5.8x came specifically from relaxing insertion-order preservation. Our shaping merges over sorted roots are that kind of constraint. Any change here must preserve digest reproducibility through deterministic partitioning.
Replace the global B-tree identifier map with a sort-based or partitioned structure. It is 6.47% of profile samples and the only component with a genuinely size-dependent constant, which fits the mild degradation we measure.
Continuous ingest throughput benchmark. Detailed below. This is not optional and should land early, because every gain won without it is unprotected.
Split the hash's two roles. Cryptographic for content-addressed naming, fast checksum for corruption detection on bytes already named. Verified practice elsewhere. A follow-on worth roughly five-fold on whatever hashing survives item 1, not a substitute for it.
Parallelism: a shared cause, not an ingest workstream
Nothing in the lifecycle is parallel. The rung runs at 0.85–0.89 of 16 cores overall (0.92 when written). Subtracting ingest's share leaves the other nine phases at 0.99. Query, reopen, recount and verification are all single-threaded too, so this cause touches 100% of lifecycle wall time and is shared with #1388. Whoever fixes it must fix it once, for the whole engine, rather than for the write path alone.
Targets
Serialized fraction of the path at or under 2%, targeting under 1%.
Effective core utilisation of at least 12 of 16 during ingest, which is what a 2% serial fraction implies. Stated explicitly so it is a number to hit rather than a derivation to perform.
The shape every comparable system converges on
Eight systems built on Arrow, Parquet and DataFusion were surveyed and all eight use the same architecture: many parallel encoders funnelling into one cheap serialized commit, with heavy investment in making that commit cheap. The design rule that follows is precise, and it is the whole of this workstream:
The serialized region contains only the metadata commit. Never encoding, sorting, hashing or copying.
Measured precedents
System
Change
Measured
GreptimeDB
Moved decoding, primary-key encoding, sorting and schema alignment out of the serialized critical section, handing the owner a pre-encoded batch
+81%, 1.20M to 2.17M points/sec, CPU 14.9 to 11.95 cores
DuckDB
Relaxed insertion-order preservation on write
5.8x on Parquet, 7.5 s to 1.3 s
DataFusion
Parallelised single-file output
6.8x, 155 s to 22.6 s, peak memory +52%
GreptimeDB's is the most directly transferable: it is the same pathology, a serialized owner holding a mutable reference while doing work that did not need to be inside the lock.
Implementation notes that will otherwise be rediscovered the hard way
The convenience Parquet writer is single-threaded encode, published at roughly 200 MiB/s per core. The parallel path is the per-column writer API, one writer per leaf column fed via channels, then appended to the row group. The async writer gives async I/O, not parallel encode; the CPU-bound work still runs inline on the calling task.
API churn: the older column-writer accessor was deprecated and then removed in a recent release. Check the version in use before following older examples.
The known trap. DataFusion's first parallel writer silently produced files missing bloom filters, column index and offset index. A content-addressed contract would not catch that, because the files are internally consistent and correctly hashed. Reproduce their production sink, not the documentation example.
Thread-pool hygiene: keep Parquet encoding off async worker threads. The reference implementations run separate I/O and CPU pools and route non-async library work through a blocking pool. One project had to fix nested work-stealing inside a large blocking pool consuming all available cores.
Ordering is the likely serialiser. DuckDB's 5.8x came specifically from relaxing insertion order. Shaping merges over sorted roots are that kind of constraint. Any change here must preserve digest reproducibility through deterministic partitioning, stable range-partition then order within partition, not by abandoning ordering. Non-deterministic parallel writes would produce different digests for the same logical graph, breaking reproducibility and dedup.
The benchmark, in detail
There is no throughput target anywhere in the project today and no continuous regression detection on ingest. The existing durable storage benchmark cases cover open, commit, recovery scan, reachability scan, garbage collection and spill compaction. None is bulk ingest. The three storage improvements that made our previous evidence stale were each validated by a bespoke one-off measurement, which is why re-running the whole ladder was the only way to know where we stood.
Requirements:
Walltime instrument, not simulation. Ingest does real durability work; a CPU simulator will not see it. That means the isolated bare-metal runner, matching how the existing durable I/O cases run.
Two or more dataset sizes, with the ratio tracked. This is the single most important design point. Our defect is that throughput degrades with size. A benchmark at one fixed size would report a flat, healthy number while that degradation continued underneath it.
Sweep a structural axis, not just total size. GreptimeDB's cardinality sweep exposed a 4.6x collapse that a pure size sweep would have blurred. Our analogous axes are edges per vertex, identifier-space density, and staged chunk count.
Report bytes read per edge as a first-class metric, alongside edges per second. That is the number that told us the problem was a constant rather than a curve.
Register it in the fail-closed benchmark inventory, so it cannot be silently deleted later.
Gate on regression, and gate on the floor. The current setup reports numbers; nothing fails when they get worse. A benchmark that cannot fail is documentation. Because the target is a floor, the benchmark must fail when throughput drops below it at any measured size, not only when it regresses relative to the previous run. A slow drift that never regresses in a single step would otherwise pass indefinitely.
Constraints
Crash recovery, corruption refusal, unsupported-format refusal and fail-closed publication outcomes are unchanged.
Existing corruption, cancellation, crash-boundary and recovery tests pass unweakened.
Digests stay reproducible for identical logical input. If parallelism or relaxed ordering changes on-disk byte layout, digests change and both reproducibility and dedup break. The fix is deterministic partitioning, not abandoning ordering.
The threat model that permits cheaper corruption detection is recorded beside the code, so a future change to it forces a revisit.
If we hand-roll parallel Parquet writing, note that DataFusion's first attempt silently produced files missing bloom filters and page indexes. A content-addressed contract would not catch that, because the files are internally consistent and correctly hashed.
Acceptance criteria
Sustained ingest is at or above 1,000,000 edges per second at every measured scale on the bench host, not only at the reference scale.
The design point clears the floor with headroom rather than meeting it exactly.
The serialized fraction of the ingest path is measured and held at or under 2%.
CPU per edge is reduced by at least 25%, from 11.93 microseconds to under 9.0. Met: 8.46 µs on f80f69fe, −29%.
Throughput does not degrade measurably across a sixteen-fold size range.
Ingest uses at least 12 of 16 available cores, and the serialized fraction is measured at or under 2%.
Bytes read per edge during ingest is reduced by at least an order of magnitude from 4,579.
A continuous ingest benchmark runs on schedule, measures at least two sizes, reports bytes per edge, is registered in the fail-closed inventory, and fails on regression.
All correctness, corruption, cancellation, crash-boundary and recovery tests pass unweakened, and digests remain reproducible.
Any change to the recorded storage or authentication design is captured in a new ADR.
Non-goals
Batching fsyncs. Traced fsync latency is 2.8 seconds against roughly 1,100 seconds of ingest, about 0.25%. Comparable systems flush every one to five seconds or never.
Redesigning compaction. Merge write amplification is already about 1.4x, better than published alternatives.
Replacing the hash primitive as the primary strategy. We already have hardware acceleration; this is worth about 6%.
Weakening durability, corruption refusal or crash safety to gain throughput.
Relationships
Contains #1384. Related to #1194, which owns storage amplification and is open only for final scale evidence. Scheduling relative to v0.6.0 is an open question.
Current plan of record (2026-09-24)
This epic owns the one-ingest multicore outcome and the direct throughput floor. #1448 owns the bounded core-scaling work and evidence; this epic is the canonical close gate for both outcomes. #1478's separate planning issue duplicated this plan and is superseded by this section; its revisions 10–14 remain archived at
docs/development/evidence/ingest-floor-plan-1478-rev14.md. GitHub's issue relationships carry execution state and prerequisites. #735 is the sole M5 closing tracker.Contract and decisions
The rate floor is 1,000,000 complete ingested edges per second at every supported scale:
T_complete_ingest(E) <= E / 1,000,000. The boundary includes registration through publication and acknowledgement. The separate multicore outcome is a material speedup for one complete ingest as more cores are made available, measured under #1448's predeclared core-use criterion on a supported durable-project host. S26 admission, a faster sub-stage, or a single fast machine does not substitute for either outcome. No macOS run is required to measure the core-count speedup; do not claim an unmeasured Apple-specific rate.Evidence and current gap
The retained S18–S22 baseline is
b6ffb088from the #1530 archive. It measured S22 at 126,099 edges/s, a 7.93× gap to the floor; rung receipts and the manifest are linked above in the retained-evidence record. A later #1481 integrated-tree checkpoint (2026-09-22,aaf20d83era) measured complete-ingest medians of 31.818 s at S18 (7.586 µs/edge; 7.59× allowance) and 547.803 s at S22 (8.163 µs/edge; 8.16× allowance). Its publication-repair A/B did not demonstrate a throughput gain. This S18/S22 checkpoint supplements, but does not replace, the archived S18–S22 ladder or the exact frozen SHA required for #1448.Sequence and ownership
The older detailed acceptance checklist below is retained for history. Under D1, historical fixed CPU-per-edge, serialized-fraction, core-count and byte ceilings are diagnostics rather than acceptance thresholds. #1448's predeclared complete-ingest core-count scaling criterion is the current multicore acceptance outcome. Close this epic on both that criterion and the rate at every supported rung, plus the no-degradation shape criterion, correctness/recovery/determinism, retained evidence, and any ADR required by a changed storage or authentication design. Keep #1476 regression ratchets active.
Any need for a process-wide admission manager, manifest format, or publication-protocol change stops the experiment that found it and routes to a separate decision under #1509. Preserve recorded UUID splitters, filesystem authority, total-memory admission, cooperative cancellation, durability, and the complete-ingest boundary. The full design cautions and reversibility notes from #1478 remain in its revision history and archived rev14 document.
Post-v0.6.0 performance owner (adjudicated 2026-10-01)
#1572 records the maintainer decision to defer the unmet 1,000,000 complete edges/s floor to post-v0.6.0. This issue remains open with its rate target and separate multicore-scaling acceptance outcome unchanged; neither is marked passed or waived. The measured S22 result was 184,749 edges/s, and the current evidence does not support expecting recent changes to close the 5.41× gap. The later eight-core S18 curve measured 1.39× one-core throughput and 1.19 effective cores, so multicore scaling is also retained as an explicit unresolved post-release outcome.
#1387 is no longer an M5 child/blocker and is no longer assigned to the M5 milestone. This issue remains the post-release owner; any future performance experiment requires a bounded issue and coordinated quiet-host allocation. v0.6.0 still requires #900/#745 S26 lifecycle certification and all other approved release gates. #1476 regression ratchets remain active.
Retained evidence (#1530 / PR #1534). The archive index distinguishes complete rung evidence from surviving controller summaries. The b6ffb088 baseline retains S18–S22 rung/result JSON, receipts, plan, projection and controller summary; its
MANIFEST.sha256digest is807609f97477e91c3cb5321261a4e9274f51573c150da88b320baf3b55c8342d. These links pin the merged archive commit from #1534. Historical raw evidence for93c041df,f80f69fe(including S24),1955f17d,ab1a713e,4fbfe84cand403fc02awas lost on 2026-09-21. Retained controller summaries support only their reported values; other historical numbers rest on the copied issue tables and comments, not missing per-operation receipts or publication-region data. The archived rev14 §12 is historical; this retained index is the current evidence location.Goal
Make GraphForge ingest fast enough that importing a graph stops being something a user plans their day around.
Floor: 1,000,000 edges per second, sustained at every supported scale.
This is a floor, not a target. Two consequences follow, and both are stronger than a target would imply.
It must hold at every scale, not at one reference point.
Our current defect is that throughput falls as the graph grows, from about 79,000 edges per second at 8 million edges to 73,000 at 67 million.Updated 2026-09-18 (f80f69fe): the shape is now flat through S22 — 101,175 / 105,415 / 105,802 / 104,390 edges/s at S18–S22 — but drops 6.6% at S24 to 97,469, with CPU/edge rising 8.46 → 9.12 µs. The drift is real and has moved to the top of the ladder. A floor forbids that shape outright. Meeting the number at one convenient size and drifting below it at a larger one is a failure, not a partial success.The design point must sit above it. You do not engineer to your floor; you engineer above it so that variance, larger inputs and future changes do not breach it.
Currently 72,989 edges per second at the reference scale, so the floor is a fourteen-fold improvement— superseded 2026-09-18. Same rung, same metric onf80f69fe: 104,390 edges/s, +43.0%. The floor is a 9.58x improvement, and the design point must be more.What guaranteeing the floor requires
Ingest spends 11.93 microseconds of CPU per edge at 67 million edges, measured on— now 8.46 µs (estimated) onfa6447ccf80f69fe, a 29% cut. It still runs on less than one of 16 cores: 0.85–0.89 effective cores, flat across a 64x edge range. CPU consumed per edge is independent of contention, so it is the stable figure to reason from; wall-clock times in this run were inflated by other work on the host. Throughput is governed by how much of the path stays serial and how much per-edge work is removed. Amdahl's law on 16 cores:Read that table as a joint constraint rather than a menu. A 5% serial fraction reaches the floor only barely, and only with a substantial efficiency gain. A 2% serial fraction clears it comfortably. Comfortable headroom across the board needs the serialized region at or under 2% and per-edge work cut by at least a quarter.
This is more achievable than the first version of this table showed, because the baseline improved. Two storage changes landed after the evidence this epic was originally written against, and they cut CPU per edge from 15.14 to 11.93 microseconds, a 21% reduction. The floor did not move; the distance to it shortened.
That turns two vague instructions into measurable budgets:
at most 2%, targeting under 1%→ at most ~10% on 16 threads (~3.7% on 8 physical cores). Relaxed 2026-09-18 because the CPU budget below is already met; the 2% figure was derived against the old baseline. Today the serial region is ~68–80%. This is the same rule the surveyed systems converge on, that the serialized region should contain only the metadata commit. No instrument computes it — see observe(storage): nothing measures the serial fraction of ingest, and #1387 budgets against it #1462.f80f69fe— 8.46 µs, a 29% cut. It did not come from removing the redundant authentication passes as predicted here; it came from fix(storage): route node-keyed families with node-only splitters, remove per-partition spill buffers (#1439) #1440 and the storage work that landed with it. perf(storage,api): remove three per-record hot loops from validate (-33% user CPU) #1458 (−33% validate user CPU, measured) is further headroom on top. The remaining distance to the floor is therefore almost entirely serial fraction.Sanity check against comparable systems
Databend is the sharpest comparison: same core count as our bench host, running at roughly twice our floor. Setting the floor below it is deliberate, since a graph edge carries identifier mapping and adjacency construction that an analytical row does not. The first draft of this epic set 400,000, anchored to the Neo4j figure, which is both a weaker peer and a result from 2016 hardware.
Why this is a product problem, not a benchmark problem
Translated into what a user waits for, at our best measured rate:
Two things make the lived experience worse than those numbers suggest.
Throughput degrades as the graph grows. Measured across a sixteen-fold range, edges per second fall about 24%. Users do not read that as a scaling curve. They read it as a tool that gets slower the more they put in it, which is the opposite of the impression a graph engine should give.
The process uses 0.85–0.89 of 16 cores (0.92 when this was written). A user watches one core pinned and fifteen idle. That reads as a broken tool regardless of the throughput number.
What this epic does and does not deliver
Ingest is 67% of lifecycle wall time, so even an infinite ingest speedup caps the total lifecycle improvement at about three-fold. That ceiling should be stated plainly, because "blazing fast ingestion" invites a broader reading than this epic delivers.
At a billion edges, with ingest at the floor and no other phase changed:
Down from roughly 7.8 hours today, a 2.7-fold lifecycle improvement from a sixteen-fold ingest improvement.
The next bottlenecks are already visible and are out of scope here: reopen proof at 0.98 hours and query at 0.72 hours, projected to a billion edges. Together they are 1.7 hours and are untouched by this work. They deserve their own epic once ingest lands.
Two caveats on the 2.71 hours. It assumes the floor holds at a billion edges, which is sixteen times beyond the reference scale and has never been measured. And it assumes the other phases stay linear, which they have been through 67 million edges but which nothing guarantees beyond it.
Baseline
Ingest is 58 to 67% of total lifecycle wall time. Measured on
fa6447cc, the ladder run completed 2026-09-17:Superseded 2026-09-18 by
clean-f80f69fe-evidence/. Ingest wall and edges are measured; effective cores is rung-wide CPU/wall and µs/edge is derived from it, so that column is estimated:At the reference scale: 265 bytes retained per edge against 4,579 bytes read for authentication, a constant factor of 17.3 that does not grow with size. Physical reads 858.3 GB, physical writes 363.8 GB, peak resident memory 196 MB.
What changed since the previous evidence
Two storage improvements landed between the earlier ladder run and this one. They are real and they are measurable at scale, though they leave the structure untouched.
The smaller rungs show no difference, so these changes target work that only appears at larger merge fan-in. Reads fell 3%, so the roughly 17-fold read amplification stands, and retained and transient storage are identical to the byte. The diagnosis below is unaffected; only the starting point moved.
Wall-clock figures for the non-ingest phases in this run ran 11% to 23% above the previous one, and whole-run core utilisation fell from 0.92 to 0.87, because other work shared the host. Those timings should not be compared across runs. Counters and per-edge figures are unaffected by contention.
Diagnosis
External research across eight comparable systems built on Arrow, Parquet and DataFusion produced a clear picture of where we differ.
We read data back in order to hash it. Nobody else does. InfluxDB 3, GreptimeDB, Iceberg, Delta, Databend, Lance, ParadeDB and SlateDB were all examined. Not one performs a read-back verification pass over data it just wrote. Where integrity exists at all it is a non-cryptographic checksum computed inline by the pass that produces the bytes, verified lazily at read time or not at all. Several hash nothing on the data path.
The serialized region contains work that does not belong there. All eight converge on the same architecture: many parallel encoders funnelling into one cheap serialized commit. The design rule that follows is that the serialized region should contain only the metadata commit, never encoding, sorting, hashing or copying. GreptimeDB measured an 81% gain from enforcing exactly that boundary. DuckDB and DataFusion each measured roughly six-fold from parallelising their write paths.
The hash primitive is not the lever. This host has the SHA-256 hardware extensions and the profile implies we already hash at about 2.65 GB/s. Switching primitives buys roughly 6% of ingest. The number of passes is the problem, not their cost per byte.
Our merge is already good. Merge write amplification is about 1.4x, better than any published log-structured design. The 21x total write amplification is therefore not the merge, and compaction redesign is not the answer.
Workstreams
Ordered by expected impact. Each becomes a child issue.
Parallelism: a shared cause, not an ingest workstream
Nothing in the lifecycle is parallel. The rung runs at 0.85–0.89 of 16 cores overall (0.92 when written). Subtracting ingest's share leaves the other nine phases at 0.99. Query, reopen, recount and verification are all single-threaded too, so this cause touches 100% of lifecycle wall time and is shared with #1388. Whoever fixes it must fix it once, for the whole engine, rather than for the write path alone.
Targets
The shape every comparable system converges on
Eight systems built on Arrow, Parquet and DataFusion were surveyed and all eight use the same architecture: many parallel encoders funnelling into one cheap serialized commit, with heavy investment in making that commit cheap. The design rule that follows is precise, and it is the whole of this workstream:
Measured precedents
GreptimeDB's is the most directly transferable: it is the same pathology, a serialized owner holding a mutable reference while doing work that did not need to be inside the lock.
Implementation notes that will otherwise be rediscovered the hard way
The benchmark, in detail
There is no throughput target anywhere in the project today and no continuous regression detection on ingest. The existing durable storage benchmark cases cover open, commit, recovery scan, reachability scan, garbage collection and spill compaction. None is bulk ingest. The three storage improvements that made our previous evidence stale were each validated by a bespoke one-off measurement, which is why re-running the whole ladder was the only way to know where we stood.
Requirements:
Constraints
Acceptance criteria
f80f69fe, −29%.Non-goals
Relationships
Contains #1384. Related to #1194, which owns storage amplification and is open only for final scale evidence. Scheduling relative to v0.6.0 is an open question.