You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
WDC Index/Arc graphs are too large for scalar add_node/add_edge loops. GraphForge’s supported path is offline convert → Arrow → chunked publish_bulk_nodes / publish_bulk_edges, but there is no checked-in tool that maps WDC integer IDs to UUIDv7 and streams batches without holding the full edge list in RAM.
Objective
Ship a focused offline ingest helper that reads Index/Arc (gzip TSV), assigns stable UUIDs, and publishes chunked bulk batches into a Parquet-backed project in exploratory mode. The tool must support capped samples (T1/T2) and be usable for larger Index/Arc tiers (T4, T5b, and Page shards at T6) when the #401 first-fail ladder reaches them.
Debt / regime
Debt type: development and data
Quality regime: A compute
Requirements
Accept paths to Index and Arc files (plain or .gz).
Stream nodes then edges in configurable Arrow batch sizes.
Map WDC integer node IDs to UUIDv7 with a deterministic or persisted mapping sufficient for reopen/replay of the same ingest.
Use public bulk APIs only (publish_bulk_nodes / publish_bulk_edges or thin Python helpers); Rust remains the engine.
Support a max-nodes / max-arcs limit for tier samples without loading the full 2012 PLD into memory.
Optionally warm CSR via index_adjacency after ingest and print disk/RSS summary hooks for spike issues.
Acceptance Criteria
Tool ingests the WDC example (106/141) end-to-end into a project directory and reopens successfully.
Tool ingests a bounded Index/Arc sample (documented tier) with peak RSS independent of full-file size (streaming).
Scenario: Example ingest
Given fetched WDC example Index/Arc files
When the ingest tool runs against a fresh project directory
Then node and edge counts match 106 / 141
And GraphForge(path) reopen yields the same counts
Scenario: Bounded sample without full-file RAM
Given 2012 PLD Index/Arc on disk
When the tool runs with a documented node/arc cap for an early tier
Then it completes without reading the entire arc file into a single in-memory list
And Cypher MATCH ()-[r]->() RETURN count(r) equals the capped arc count
Do not implement BVGraph/Pajek parsers here; accept Index/Arc only (2014 PLD/Host edges require a prior conversion step owned elsewhere or documented externally for test(scale): WDC T0→T6 first-fail escalation spike #401).
Observability
Log batch sizes, rows published, elapsed, and peak RSS if available. No UUID property dumps.
Security And Privacy
Local filesystem only; no network in the ingest tool.
Testing
Unit tests for TSV parsing / ID mapping on the example files.
Integration test gated or marked heavy for sample ingest if CI-appropriate; otherwise scripted manual evidence for larger tiers.
Documentation
Update docs/guide/datasets/wdc-hyperlink-graph.md ingest section with exact commands once the tool lands.
Non-Goals
BVGraph/Pajek native parsers (conversion stays out of tree for this issue)
Problem
WDC Index/Arc graphs are too large for scalar
add_node/add_edgeloops. GraphForge’s supported path is offline convert → Arrow → chunkedpublish_bulk_nodes/publish_bulk_edges, but there is no checked-in tool that maps WDC integer IDs to UUIDv7 and streams batches without holding the full edge list in RAM.Objective
Ship a focused offline ingest helper that reads Index/Arc (gzip TSV), assigns stable UUIDs, and publishes chunked bulk batches into a Parquet-backed project in exploratory mode. The tool must support capped samples (T1/T2) and be usable for larger Index/Arc tiers (T4, T5b, and Page shards at T6) when the #401 first-fail ladder reaches them.
Debt / regime
Requirements
.gz).publish_bulk_nodes/publish_bulk_edgesor thin Python helpers); Rust remains the engine.index_adjacencyafter ingest and print disk/RSS summary hooks for spike issues.Acceptance Criteria
BDD Completion Scenarios
Scenario: Example ingest
Given fetched WDC example Index/Arc files
When the ingest tool runs against a fresh project directory
Then node and edge counts match 106 / 141
And
GraphForge(path)reopen yields the same countsScenario: Bounded sample without full-file RAM
Given 2012 PLD Index/Arc on disk
When the tool runs with a documented node/arc cap for an early tier
Then it completes without reading the entire arc file into a single in-memory list
And Cypher
MATCH ()-[r]->() RETURN count(r)equals the capped arc countImplementation Notes
Observability
Security And Privacy
Testing
Documentation
docs/guide/datasets/wdc-hyperlink-graph.mdingest section with exact commands once the tool lands.Non-Goals
graphforge.datasetscatalog APIRelated Issues
docs/guide/graph-construction.mdOpen Questions