Skip to content

perf(api): import normalization is 13% of validate at exactly one effective core #1472

Description

@DecisionNerd

Execution prerequisite (#1478 rev 13): blocked by #1477 in native issue relationships. This issue owns the normalization/input-overlap experiment that precedes #1448's decomposition. Its result must be evaluated before that dependent implementation; success is not presumed.

Summary

normalize_import_node_chunk / normalize_import_edge_chunk are 13.0% of import-session validate at S18 rung scale, running at 1.01 effective cores — perfectly serial, on a 16-thread host.

Measured on real S18 Graph500 data, release build, quiet host, validate wall 36.14 s:

stage wall % of validate eff. cores
shaping 19.88 s 55.0% 0.93
canonical_encoding 6.50 s 18.0% 0.88
normalization 4.69 s 13.0% 1.01
append 4.23 s 11.7% 0.87
parquet decode + residual ~0.84 s 2.3% —

97.7% of validate accounted. Normalization is the third-largest stage, larger than every region inside shaping except after_chain (16.5%) and chunk_loop (14.7%).

Why this one is worth doing

Nobody had measured it. Every previous analysis of the ingest path — including my own on #1464 — divided up shaping and treated the rest as overhead. It is not overhead; it is a stage.

And unlike most of the serial path, this work has no ordering dependency. assign_surrogates is sequential because surrogates are a running counter in global order; finish_optional's consume is sequential because a linear digest authenticates the output. Normalization is per-batch over independent input chunks, and it is serial because nothing made it otherwise, not because it must be.

Where the work is

crates/graphforge-api/src/bulk_construction/normalization.rs:1305, called per batch from the staging loop in import_session.rs. For edges it builds a BulkNodeRow per candidate endpoint — roughly 8.4 million at S18 (4.19M edges x 2) — before normalize_bulk_edges runs.

Two independent axes, either of which is worth measuring alone:

  1. Within a batch. Row-wise validation over a RecordBatch with no cross-row dependency beyond within-chunk duplicate detection.
  2. Across batches. epic(storage): stop maintaining a hand-rolled dataflow engine; reuse Arrow/DataFusion where determinism and durability allow #1456's implementation item 3 already proposes overlapping input preparation with append, "preserving batch order, replay identifiers and manifest semantics", with the queue governed by bytes rather than batch count. Perfect overlap of normalization with append caps at min(4.69, 4.23) — about 11.6% of validate — so the two axes compose rather than substitute.

Acceptance

Context that should temper expectations

This is one stage of nine, and the whole path measures 0.92 effective cores with nothing above 1.16. Fixing normalization alone moves validate by at most ~13%, and #1387's floor is 9.2x away. It is worth doing because it is the cleanest available parallelism in the path and it is measured, not because it closes the gap.

Related: #1387, #1456 (implementation item 3), #1464 (where the rung-scale attribution is recorded).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions