Problem
Confirmed on current main at 6dc51fa47a9658d26a3f6e51e8a8261aaf9b4b86.
Each variable-length path-node UDF invocation calls hydrate_path_node_children at crates/graphforge-rel/src/expr.rs:11998-12000. The helper reads every node batch and builds a graph-wide UUID map at crates/graphforge-rel/src/expr.rs:12057-12083, then reads and concatenates every property stem and indexes every row at crates/graphforge-rel/src/expr.rs:12111-12166.
Work is proportional to total graph/property volume rather than requested nodes and repeats across UDF/output batches. A small nodes(p) result can scan and materialize the whole graph and recreate contiguous Arrow limits. Closed M4 issue #341 addressed analyst-output shaping, not this Cypher UDF.
Objective
Hydrate variable-length path nodes demand-first and batchwise, with work and peak memory bounded by requested UUIDs plus configured resource limits.
Requirements
- Deduplicate requested UUIDs before reads.
- Push identity selection into batchwise node/property reads or an indexed gather path; do not concatenate complete tables.
- Avoid rebuilding graph-wide UUID maps per UDF/output batch.
- Preserve full labels, properties, ordering, nullability, repeated-node semantics, cancellation, and structured resource errors.
- Consume the established embedded resource/batch policy; no independent pool or unbounded cache.
- Keep Rust authoritative and bindings thin.
Acceptance Criteria
BDD Completion Scenarios
Scenario: A small path does not materialize a large graph
Given a graph with many nodes and property stems
And a path containing a small selected subset
When nodes(p) is evaluated
Then GraphForge gathers only needed UUIDs within the resource policy
And returns complete canonical node values.
Scenario: Batch boundaries do not alter hydration
Given the same graph/path under different accepted batch sizes
When path nodes are hydrated
Then labels, properties, ordering, repeats, nulls, and errors are identical
And only bounded I/O/allocation observations differ.
Implementation Notes
Likely surfaces: crates/graphforge-rel/src/expr.rs, storage/catalog readers, UUID lookup/index facilities, resource policy, and path-expression tests.
Issue #705 must land first so the optimization preserves the authoritative full-label contract. A native blocked-by relationship will encode this.
Observability
Use aggregate test/benchmark counters for rows/batches scanned, rows gathered, bytes materialized, and peak memory. Never log values, UUIDs, properties, or paths.
Security And Privacy
Any cache/index must be generation-safe, bounded, instance-scoped, and free of user-data logging.
Testing
Use a sparse-selection fixture with irrelevant nodes and multiple property stems. Cover repeated/missing UUIDs, cancellation, reopen, rollover, and deterministic results. Compare structural counters and peak memory, not timing alone. Preserve #705's full-label regression.
Documentation
Update execution/scale docs if hydration gains a documented resource/indexing contract.
Non-Goals
Related Issues
Problem
Confirmed on current
mainat6dc51fa47a9658d26a3f6e51e8a8261aaf9b4b86.Each variable-length path-node UDF invocation calls
hydrate_path_node_childrenatcrates/graphforge-rel/src/expr.rs:11998-12000. The helper reads every node batch and builds a graph-wide UUID map atcrates/graphforge-rel/src/expr.rs:12057-12083, then reads and concatenates every property stem and indexes every row atcrates/graphforge-rel/src/expr.rs:12111-12166.Work is proportional to total graph/property volume rather than requested nodes and repeats across UDF/output batches. A small
nodes(p)result can scan and materialize the whole graph and recreate contiguous Arrow limits. Closed M4 issue #341 addressed analyst-output shaping, not this Cypher UDF.Objective
Hydrate variable-length path nodes demand-first and batchwise, with work and peak memory bounded by requested UUIDs plus configured resource limits.
Requirements
Acceptance Criteria
concat_batchessolely for path hydration.BDD Completion Scenarios
Scenario: A small path does not materialize a large graph
Given a graph with many nodes and property stems
And a path containing a small selected subset
When
nodes(p)is evaluatedThen GraphForge gathers only needed UUIDs within the resource policy
And returns complete canonical node values.
Scenario: Batch boundaries do not alter hydration
Given the same graph/path under different accepted batch sizes
When path nodes are hydrated
Then labels, properties, ordering, repeats, nulls, and errors are identical
And only bounded I/O/allocation observations differ.
Implementation Notes
Likely surfaces:
crates/graphforge-rel/src/expr.rs, storage/catalog readers, UUID lookup/index facilities, resource policy, and path-expression tests.Issue #705 must land first so the optimization preserves the authoritative full-label contract. A native blocked-by relationship will encode this.
Observability
Use aggregate test/benchmark counters for rows/batches scanned, rows gathered, bytes materialized, and peak memory. Never log values, UUIDs, properties, or paths.
Security And Privacy
Any cache/index must be generation-safe, bounded, instance-scoped, and free of user-data logging.
Testing
Use a sparse-selection fixture with irrelevant nodes and multiple property stems. Cover repeated/missing UUIDs, cancellation, reopen, rollover, and deterministic results. Compare structural counters and peak memory, not timing alone. Preserve #705's full-label regression.
Documentation
Update execution/scale docs if hydration gains a documented resource/indexing contract.
Non-Goals
Related Issues