Skip to content

perf(storage): make one complete ingest scale with available cores #1448

Description

@DecisionNerd

Current architecture and child stack (2026-09-27)

The #1599 shaping and #1600 encoding changes are merged. The latest S18 eight-core curve records 24.67 s complete ingest, 1.39× one-core throughput and 1.19 effective cores; canonical encoding reaches 2.1086 CPU/wall. Routing (4.18 s, 0.89 CPU/wall) and append (4.19 s, 0.88 CPU/wall) remain bounded serial-work candidates. Measured evidence.

Order Native child State / scope
Complete #1599 Shaping partition filesystem work; merged
Complete #1600 Canonical topology and adjacency encoding; merged
Next #1606 Route independent staged families on admission lanes; preserve shared progress/retirement boundary
After #1606 #1607 Prepare one construction chunk on admission lanes; preserve single durable intent/receipt owner

Both open children natively block this issue; #1607 is natively blocked by #1606. The dependency freezes routing's measured disposition before append's baseline/implementation, rather than creating a code dependency between the mechanisms. One concern per PR. Do not start dependent implementation before its prerequisite merges; a documented no-go can satisfy the prerequisite without production adoption.

Architecture decision (quality regime A). Extend existing leased-lane seams with exclusive task ownership and ordered coordinator reduction. Routing workers own disjoint family partitioners for one receipt; the coordinator retains seal/progress/retirement. Append workers prepare one bounded chunk; the coordinator retains intent, durable installs, receipt chain and checkpoint advancement. Preserve cancellation, authentication, duplicate decisions, actual resource accounting and publication-boundary determinism. This follows ADRs 0038/0045/0046/0047; no new ADR, public API, durable format or process-wide scheduler is needed.

Options considered: (1) retain serial behavior with a measured no-go, the lowest-risk outcome when benefit is insufficient; (2) bounded independent work under existing coordinators, the selected experimental path with limited migration cost; (3) cross-receipt/cross-chunk asynchronous writes, rejected here because they expand recovery authority, buffering and ordering risk. The chosen path is reversible before adoption; a protocol/format/admission-policy change requires a new decision rather than scope growth.

Finite experiment limit. With the measured 34.17 s one-core and 24.67 s eight-core times held fixed, ideal four-way acceleration of both 4.18 s routing and 4.19 s append regions gives about 18.39 s, or 1.86× throughput. Even removing both regions entirely gives only about 2.10×. These are optimistic ceilings, not forecasts: task balance, I/O, overlap and extra work can reduce the gain, and the one-core path must be remeasured. Do not promise that these two children achieve the parent's 2× throughput and 2.0 CPU/wall criteria.

Each child records its performance gate before candidate measurements and either merges a qualifying implementation or records a bounded no-go. After these outcomes, this canonical issue owns the combined core-use decision, S18/S20/S22 reconciled region evidence, required changed-surface correctness/recovery evidence, the measured disposition on #1387 and input to #1572. If a release-sized route remains unsupported, record the limit for #1572 instead of automatically creating another experiment stack. Closing the children alone does not close #1448.


Scope of record, 2026-09-24. This issue owns useful use of the cores available to one complete ingest. Partition-level shaping is the first bounded implementation hypothesis, chosen from the #1509 reuse decision. The #1572 release checkpoint consumes this issue's core-scaling evidence; the 1M edges/s floor remains #1387's separate target. Earlier shape-only wording is superseded by the scope and acceptance below.

Summary

One complete ingest still uses roughly one core through its largest regions. Shaping is the first measured parallelism lever, with encoding and other regions to be reassessed after the shaping experiment. On the current tree (#1439 candidate 6d6b7f17, quiet S19, 2026-09-20, import_command/validate region diagnostics from the #1456 close evidence):

Region Wall s Process CPU s CPU / wall Share of validate
validate (whole) 50.08 48.66 0.97 100%
seal 39.48 36.00 0.91 79%
seal / shaping 23.88 21.39 0.90 48%
seal / canonical_encoding 15.51 14.52 0.94 31%
append 7.83 6.80 0.87 16%
normalization (bounded parallel, #1472) 1.90 5.53 2.91 4%

Shaping and encoding together are 79% of validate and both sit at 0.90–0.94 effective cores on a 16-thread host. Every remaining lever for the #1387 floor that is not byte removal runs through this region.

The historical S20 figure in earlier revisions of this issue ("seal is 113.2 s of a 165.7 s ingest, one call, single-threaded") measured the API seal timer, which wraps shape and encode. Any claim about shaping must use the seal/shaping region, not the operation timer.

What the spikes established

Scope

Make one complete ingest use available cores efficiently on an admitted durable-project host, while preserving its published results and resource bounds. First test partition-level shaping, using the retain/adopt/hybrid facility selected by #1509:

After the shaping A/B, measure the complete-ingest core-count curve again and identify the new limiting region. If a credible, independently reviewable second mechanism is required for useful scaling, create a bounded native sub-issue that blocks this canonical issue; do not add another concern to the shaping PR. If evidence rules out a release-sized route, record the limit and feed #1572's decision instead of running an unbounded experiment series.

The floor, durability, publication-boundary determinism (ADR 0038), cancellation, and complete-ingest timing boundary are unchanged. An isolated shaping-region speedup or a 10% whole-ingest gain does not by itself satisfy this issue's multicore objective.

Predeclared criterion

Shaping adoption gate. Compare the candidate with its integrated baseline under equal resources: three alternating quiet S18 pairs and one S20 pair (protocol in docs/development/construction-reuse-inventory-protocol-1505.md §5.3; BenchExec, CPUs 0–15, 4 GiB). Report seal/shaping wall, CPU and CPU/wall; whole-ingest wall and CPU; and reconciled region residuals. Adopt only if whole-ingest median wall improves by at least 10% at both S18 and S20, with identical shaped and encoded payload digests. Below that, retain the baseline and record the no-go. This threshold is a noise/benefit gate for the shaping change, not a core-use or 1M SLO threshold.

Core-use gate. Before judging the candidate, record a fixed-input complete-ingest scale-up curve on the same admitted durable-project host at 1, 2, 4, 8 and maximum available logical CPUs (omit duplicate counts). Set the worker limit and CPU allocation consistently; report actual usable CPUs, wall, CPU-seconds, effective cores (CPU/wall), throughput relative to one core, and limiting phase/region at each point. Use one representative S18-scale input and repeat only points whose noise would alter the decision. Repeat the relevant curve after the adopted change. Explain CPU starvation, I/O wait, memory pressure, scheduler limits and workload size before calling a low CPU/wall ratio serial code. Record a maintainer-agreed material scale-up criterion after the baseline curve and before candidate results; do not choose it retrospectively. A supported multicore Linux host is sufficient for this scale-up gate; no macOS or Apple-specific benchmark is required.

Acceptance

  • A baseline complete-ingest core-count curve, plus a candidate curve if a change is adopted, shows actual CPU availability, speedup, effective cores and the limiting region. The material scale-up criterion is recorded before candidate results; the evidence does not claim an unmeasured Apple-specific rate.
  • Region diagnostics before and after at S18, S20 and S22 report seal/shaping and whole-ingest wall, CPU and CPU/wall, with residuals reconciled to the root. Scheduler counters are reported where available.
  • The shaping change passes the predeclared whole-ingest A/B gate or receives a measured no-go with no production adoption.
  • After shaping, either the declared core-use criterion is met on the admitted host or the remaining bottleneck and bounded next mechanism are documented. Credible independent work becomes a native blocker; an evidence-backed no-go goes to decision(storage): adjudicate the 1M ingest SLO for v0.6.0 and shared-host time #1572. Do not mark one-core ingest as good multicore use.
  • Durable artifacts remain self-consistent through the determinism suite at every recorded partition count and across interrupted/resumed import. Evidence, cancellation and cleanup stay order-independent under concurrent lanes, including a deliberately permuted schedule test when lanes land.
  • TCK, GDC, storage release, corruption-refusal and recovery checks pass for adopted code. Record the measured disposition on epic(storage): scale complete ingest across cores and reach 1M edges/s #1387 and supply its evidence to decision(storage): adjudicate the 1M ingest SLO for v0.6.0 and shared-host time #1572.

Non-goals

The refuted family-axis and surrogate-prefix-sum proposals; removing durability barriers; wholesale engine migration (owned by #1504/#1509); publication redesign (owned by #1481); asserting the 1M floor from this issue alone. Unbounded full-ladder repetitions, lowering correctness guarantees or moving work outside complete ingest are also out of scope.

Relationships

Native sub-issue and blocker of #1387; #1572 is blocked by this issue and owns the release decision. #1509 and the earlier mechanism/baseline prerequisites are closed. The active native blockers and sequence are #1606 then #1607, as recorded above. #1433 independently blocks #1572. Historical coordination was in #1478, now superseded by #1387's plan of record. Related: #1459 (bounded partition window), #1464 (refuted chain proposal), #1504 (reuse evaluation), #1481 (publication).

Hypothesis, baseline freeze and boundaries (2026-09-21, #1478 rev 14)

Hypothesis under test. Moving a measured, row-proportional shaping operation from ordered coordinator consumption into independent partition tasks reduces complete-ingest wall enough to pass the predeclared 10% gate. Choose the operation from the integrated region profile, not from the historical tables above. This is the first bounded structural hypothesis for the multicore objective. The explicit-exchange, owned-artifact pipeline proposed and critiqued on 2026-09-21 is recorded as a hypothesis row in #1509 and is not built here.

Baseline freeze rule. Record the integrated baseline SHA and binary digest when the A/B begins and rebase the candidate onto that SHA; the adopt decision is measured against it. Changes landing on main after the freeze do not invalidate the A/B unless they touch seal/shaping, publication or the receipt's boundaries; then rerun the pairs on the new base before adoption. Copy evidence to a retained location (#1530) before another ladder or test suite touches the ladder root.

Prerequisite state. The baseline tree contains #1452 and the #1526 fix. Reuse the retained b6ffb08 S18–S22 ladder as historical context, then freeze the integrated current-main SHA and binary digest for the A/B and core-count curves. Do not use lost ab1a713 per-operation receipts as current-tree evidence.

Boundaries. No verification check is removed or relocated; moving work to verify or outside the ingest timer does not count as improvement. Endpoint resolution stays as is unless evidence justifies a change. Existing admission, publication and recovery contracts stay as they are. Recorded UUID splitters remain the durable partition authority; execution subdivision cannot separate records that share one duplicate decision. If the work turns out to need a process-wide admission manager, a manifest format or a publication-protocol change, stop, record the dependency and its migration cost, and route it to #1509 for a separate decision.

Ceiling. Shaping is 48% of validate at ~0.90 effective cores and validate is less than all of ingest, so eliminating shaping's cost entirely bounds the gain below 1.92×. The remaining S22 gap on ab1a713e is 6.45×. Shaping alone cannot meet the floor; this issue now owns the multicore outcome through bounded follow-up or a measured no-go, while #1387 retains the direct rate gate.

Continuation and revisit triggers. After the A/B, recompute the remaining rate gap and core-count scaling from measured complete ingest before any further structural work. Before production adoption, reopen the decision if: incremental import into an existing generation would require full adjacency reconstruction under the lane design; the design needs a query-admission contract change (ADR 0026); or substantial prepared work is discarded after publication conflicts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions