Skip to content

feat(io): chunked WDC Index/Arc Arrow bulk ingest tool #400

Description

@DecisionNerd

Problem

WDC Index/Arc graphs are too large for scalar add_node/add_edge loops. GraphForge’s supported path is offline convert → Arrow → chunked publish_bulk_nodes / publish_bulk_edges, but there is no checked-in tool that maps WDC integer IDs to UUIDv7 and streams batches without holding the full edge list in RAM.

Objective

Ship a focused offline ingest helper that reads Index/Arc (gzip TSV), assigns stable UUIDs, and publishes chunked bulk batches into a Parquet-backed project in exploratory mode. The tool must support capped samples (T1/T2) and be usable for larger Index/Arc tiers (T4, T5b, and Page shards at T6) when the #401 first-fail ladder reaches them.

Debt / regime

  • Debt type: development and data
  • Quality regime: A compute

Requirements

  1. Accept paths to Index and Arc files (plain or .gz).
  2. Stream nodes then edges in configurable Arrow batch sizes.
  3. Map WDC integer node IDs to UUIDv7 with a deterministic or persisted mapping sufficient for reopen/replay of the same ingest.
  4. Use public bulk APIs only (publish_bulk_nodes / publish_bulk_edges or thin Python helpers); Rust remains the engine.
  5. Support a max-nodes / max-arcs limit for tier samples without loading the full 2012 PLD into memory.
  6. Optionally warm CSR via index_adjacency after ingest and print disk/RSS summary hooks for spike issues.

Acceptance Criteria

  • Tool ingests the WDC example (106/141) end-to-end into a project directory and reopens successfully.
  • Tool ingests a bounded Index/Arc sample (documented tier) with peak RSS independent of full-file size (streaming).
  • Deterministic operation UUIDs / idempotent retry behavior documented.
  • Failure modes fail closed (bad TSV, missing endpoints, UUID conflicts) without partial silent success.

BDD Completion Scenarios

Scenario: Example ingest
Given fetched WDC example Index/Arc files
When the ingest tool runs against a fresh project directory
Then node and edge counts match 106 / 141
And GraphForge(path) reopen yields the same counts

Scenario: Bounded sample without full-file RAM
Given 2012 PLD Index/Arc on disk
When the tool runs with a documented node/arc cap for an early tier
Then it completes without reading the entire arc file into a single in-memory list
And Cypher MATCH ()-[r]->() RETURN count(r) equals the capped arc count

Implementation Notes

Observability

  • Log batch sizes, rows published, elapsed, and peak RSS if available. No UUID property dumps.

Security And Privacy

  • Local filesystem only; no network in the ingest tool.

Testing

  • Unit tests for TSV parsing / ID mapping on the example files.
  • Integration test gated or marked heavy for sample ingest if CI-appropriate; otherwise scripted manual evidence for larger tiers.

Documentation

  • Update docs/guide/datasets/wdc-hyperlink-graph.md ingest section with exact commands once the tool lands.

Non-Goals

  • BVGraph/Pajek native parsers (conversion stays out of tree for this issue)
  • graphforge.datasets catalog API
  • Analyst-verb correctness beyond reopen + LIMIT Cypher smoke
  • Owning the test(scale): WDC T0→T6 first-fail escalation spike #401 ladder run itself (this tool enables tiers; spike owns first-fail execution)

Related Issues

Open Questions

  • None

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requesttoolingDeveloper tooling and automation

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions