You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Produce the inventory and version-grounded comparison protocol required by #1504 before any replacement prototype is measured. This issue is the sequencing gate for the mechanism spikes.
Scope
Map every custom sorting, partitioning, spilling, scheduling, buffering, cancellation, and memory-accounting mechanism in the construction path to callers, invariants, tests, and maintenance burden. Separate existing library reuse from candidate new reuse.
Record current-main baseline revision and exact dependency versions (arrow, datafusion, tokio, Rayon, related crates). Link version-matched upstream docs for capability and limitation claims.
Build a mechanism-by-candidate matrix: current implementation, applicable DataFusion facilities, Arrow kernels/batch facilities, Tokio coordination facilities, and meaningful hybrids. Include Rayon as the scheduling baseline. Do not assume each library replaces every mechanism.
Separate execution partitioning from durable layout, generic mechanisms from graph semantics, transient spill from recovery authority, and library-accounted memory from total admitted memory.
Before any spike measures results, record workloads, resource envelopes, repetitions, uncertainty treatment, regression thresholds, correctness gates, and how maintenance benefits will be weighed against performance costs. Do not invent a universal speedup threshold. Explicitly note that any production tradeoff changing an existing requirement needs maintainer acceptance.
Surfaces
Likely: crates/graphforge-storage/src/graph_construction*, construction_directory.rs, construction determinism/lifecycle tests, API import/resource-policy code. Related evidence from #1456 / #1465 may be cited with stated limitations.
Acceptance criteria
Inventory covers sorting, partitioning, spilling, and scheduling, plus memory/cancellation interactions.
Candidate matrix is version-grounded and distinguishes applicable vs inapplicable facilities per mechanism.
Comparison protocol is recorded in-repo before spike benchmark results land; later protocol changes are explained, not retrofitted.
Workloads (empty/tiny, normal, duplicate-heavy, skewed/hub, variable-width/nested properties, spill-forcing) and S18–S20 quiet-host measurement expectations are named without claiming S26 qualification.
Correctness gates (identities, duplicate policy, surrogate/endpoint mappings, properties, ADR 0038 publication where required, CAS/receipts, structured errors, atomic publication) are listed as hard gates for later spikes.
Purpose
Produce the inventory and version-grounded comparison protocol required by #1504 before any replacement prototype is measured. This issue is the sequencing gate for the mechanism spikes.
Scope
mainbaseline revision and exact dependency versions (arrow,datafusion,tokio, Rayon, related crates). Link version-matched upstream docs for capability and limitation claims.Surfaces
Likely:
crates/graphforge-storage/src/graph_construction*,construction_directory.rs, construction determinism/lifecycle tests, API import/resource-policy code. Related evidence from #1456 / #1465 may be cited with stated limitations.Acceptance criteria
Non-goals
Runnable replacement prototypes; performance claims; ADR conclusions; production migration.
Parent
Native sub-issue and blocker of #1504.