Skip to content

test(scale): define the billion-live-edge public-facade contract and first-fail ladder #736

Description

@DecisionNerd

Problem

The existing ignored SCALE20 Graph500 facade runner is useful but retains all raw tuples in memory, and historical large-scale evidence does not define a reproducible billion-live-edge product contract. SCALE26/edgefactor-16 produces 1,073,741,824 raw attempts, but self-loop and duplicate policy can reduce the live persisted count below one billion.

Objective

Define a deterministic, public-facade scale contract and preflight ladder that measures the first real bottleneck before the final certification.

Requirements

  • Specify rungs from the current SCALE20 baseline through intermediate sizes to SCALE26.
  • Replace all-in-memory raw tuple retention with bounded-memory generation/spill suitable for the ladder.
  • Define vertex/edge identity, self-loop, duplicate, directionality, and deterministic seed policy.
  • Measure raw attempts, accepted/rejected rows, unique live persisted edges, elapsed time, peak RSS, and peak disk.
  • Stop at and report the first envelope violation.
  • Keep execution opt-in/ignored locally and runnable as an explicitly provisioned CI or release evidence job.

Acceptance criteria

  • A versioned scale profile defines the ladder, seed, policies, envelope, metrics, and exact invocation.
  • The runner uses only public GraphForge surfaces for graph construction and verification.
  • Generator memory is bounded independently of total edge count.
  • Tests prove raw attempts cannot be mistaken for live persisted edges.
  • At least one practical lower rung runs in normal validation; larger rungs are safely opt-in.
  • Documentation explains how to reproduce and interpret first-fail evidence.

Scenarios

  • Given duplicate and self-loop inputs, when the profile runs, then raw, rejected, and live counts are distinct and reconcile.
  • Given a configured RSS or disk ceiling is crossed, when the next rung would start, then the runner stops with the first failing phase and measured resources.
  • Given the same seed and profile, when the bounded generator runs twice, then its canonical input fingerprint and counts match.

Testing and observability

Add unit tests for count reconciliation and deterministic generation, an integration test for a small profile, and phase-level structured metrics suitable for final certification.

Non-goals

This issue does not itself certify one billion live edges or redesign storage.

Related issues

Canonical tracker: #735. Historical context: #338, #408, #410, #710.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    coreCore source code changestestingTest coverage and testing infrastructure

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions