You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
perf(api): import normalization is 13% of validate at exactly one effective core #1472
Execution prerequisite (#1478 rev 13): blocked by #1477 in native issue relationships. This issue owns the normalization/input-overlap experiment that precedes #1448's decomposition. Its result must be evaluated before that dependent implementation; success is not presumed.
Summary
normalize_import_node_chunk / normalize_import_edge_chunk are 13.0% of import-session validate at S18 rung scale, running at 1.01 effective cores — perfectly serial, on a 16-thread host.
Measured on real S18 Graph500 data, release build, quiet host, validate wall 36.14 s:
stage
wall
% of validate
eff. cores
shaping
19.88 s
55.0%
0.93
canonical_encoding
6.50 s
18.0%
0.88
normalization
4.69 s
13.0%
1.01
append
4.23 s
11.7%
0.87
parquet decode + residual
~0.84 s
2.3%
—
97.7% of validate accounted. Normalization is the third-largest stage, larger than every region inside shaping except after_chain (16.5%) and chunk_loop (14.7%).
Why this one is worth doing
Nobody had measured it. Every previous analysis of the ingest path — including my own on #1464 — divided up shaping and treated the rest as overhead. It is not overhead; it is a stage.
And unlike most of the serial path, this work has no ordering dependency.assign_surrogates is sequential because surrogates are a running counter in global order; finish_optional's consume is sequential because a linear digest authenticates the output. Normalization is per-batch over independent input chunks, and it is serial because nothing made it otherwise, not because it must be.
Where the work is
crates/graphforge-api/src/bulk_construction/normalization.rs:1305, called per batch from the staging loop in import_session.rs. For edges it builds a BulkNodeRow per candidate endpoint — roughly 8.4 million at S18 (4.19M edges x 2) — before normalize_bulk_edges runs.
Two independent axes, either of which is worth measuring alone:
Within a batch. Row-wise validation over a RecordBatch with no cross-row dependency beyond within-chunk duplicate detection.
validate wall falls, measured on a quiet host across at least three runs.
Output identical: the four fixed-width run digests and every published artifact byte-for-byte unchanged, per ADR 0038's boundary.
Within-chunk duplicate detection and every existing validation refusal preserved — this is a parallelisation, not a relaxation.
Batch order, replay identifiers and manifest semantics unchanged, so recovery and resume behave identically.
Context that should temper expectations
This is one stage of nine, and the whole path measures 0.92 effective cores with nothing above 1.16. Fixing normalization alone moves validate by at most ~13%, and #1387's floor is 9.2x away. It is worth doing because it is the cleanest available parallelism in the path and it is measured, not because it closes the gap.
Related: #1387, #1456 (implementation item 3), #1464 (where the rung-scale attribution is recorded).
Summary
normalize_import_node_chunk/normalize_import_edge_chunkare 13.0% ofimport-session validateat S18 rung scale, running at 1.01 effective cores — perfectly serial, on a 16-thread host.Measured on real S18 Graph500 data, release build, quiet host,
validatewall 36.14 s:shapingcanonical_encodingappend97.7% of
validateaccounted. Normalization is the third-largest stage, larger than every region inside shaping exceptafter_chain(16.5%) andchunk_loop(14.7%).Why this one is worth doing
Nobody had measured it. Every previous analysis of the ingest path — including my own on #1464 — divided up shaping and treated the rest as overhead. It is not overhead; it is a stage.
And unlike most of the serial path, this work has no ordering dependency.
assign_surrogatesis sequential because surrogates are a running counter in global order;finish_optional's consume is sequential because a linear digest authenticates the output. Normalization is per-batch over independent input chunks, and it is serial because nothing made it otherwise, not because it must be.Where the work is
crates/graphforge-api/src/bulk_construction/normalization.rs:1305, called per batch from the staging loop inimport_session.rs. For edges it builds aBulkNodeRowper candidate endpoint — roughly 8.4 million at S18 (4.19M edges x 2) — beforenormalize_bulk_edgesruns.Two independent axes, either of which is worth measuring alone:
RecordBatchwith no cross-row dependency beyond within-chunk duplicate detection.min(4.69, 4.23)— about 11.6% of validate — so the two axes compose rather than substitute.Acceptance
validatewall falls, measured on a quiet host across at least three runs.Context that should temper expectations
This is one stage of nine, and the whole path measures 0.92 effective cores with nothing above 1.16. Fixing normalization alone moves
validateby at most ~13%, and #1387's floor is 9.2x away. It is worth doing because it is the cleanest available parallelism in the path and it is measured, not because it closes the gap.Related: #1387, #1456 (implementation item 3), #1464 (where the rung-scale attribution is recorded).