Skip to content

test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900

Description

@DecisionNerd

Where the current numbers live

This issue deliberately records no free-space, available-capacity, spare-after-reserve or deficit figures. Every such number sampled the host at a moment and was stale within days. Three superseded snapshots previously accumulated here and contradicted each other. Read the current values from the sources below instead, at the time you need them.

Live admission decision and measured capacity. The controller measures actual work-root free space and projects the next rung from completed adjacent rungs. A dry run decides without executing or mutating anything:

PYTHONPATH=benchmarks/harness python3 -m graphforge_bench.progressive_host_run \
  --dry-run --rung S24 \
  --work-root <work root> --output-dir <fresh empty dir> \
  --gf <gf> --certify <certify> --generator <generator>

It prints the plan with native_capacity.free_bytes, native_capacity.reserved_headroom_bytes, the full projected block and the decision. A dry run refuses to reuse a directory holding a prior attempt, so always point --output-dir at a fresh path.

Per-rung evidence from the last ladder attempt. On OVHC-AGENCY, /home/ubuntu/graphforge-ladder/1194-710c6c64-evidence/:

  • s<N>-projection.json — the admission decision for rung N, its checks map, limits, and projected demand
  • s<N>-rung.json — measured live_edges, metrics.peak_rss_bytes, metrics.wall_seconds, storage attribution
  • s<N>-result.json — pass/fail, typed failure, and frozen identities
  • frozen-executables.json — the exact binaries the attempt ran

Reserve policy. The reserve is an input, not a constant. Set it from the work root at launch: at least 15% of measured available bytes, rounded up. A reserve carried forward from an older sample is not automatically correct for today's filesystem.

Do not treat any file under docs/ as the capacity ledger. Those files are point-in-time reports and go stale the same way.

Current status — 2026-09-17

The ladder has executed S18 → S19 → S20 → S22 on merged source 710c6c64f4718c0664d08bdd3aafe21de4a29eba, with native admission before each successor and all ten ordinary lifecycle phases passing at every rung. Accepted rung workspaces were reclaimed before the next rung. S22 holds 4,194,304 nodes and 67,108,864 live edges.

S24, S25 and S26 have never been executed or authorized. That is the entire remaining gap for this issue, #1194 and #745.

Measured resource evidence, commit 710c6c64

Rung Live edges Peak process RSS (bytes) Wall seconds
S18 4,194,304 183,107,584 88
S19 8,388,608 189,415,424 176
S20 16,777,216 199,135,232 360
S22 67,108,864 262,049,792 1,683

Both recorded projections return admitted with rss_bounded_or_plateaued: true. The declared per-rung ceiling is 4 GiB; S22 peaked at roughly 6% of it.

Premortem for the S26 run — 2026-09-17

Correction to the prior premortem

An earlier premortem in this session concluded that S24 admission was still refused on RSS growth and that the plan was capped at S22. Both claims were wrong. They were read from the dated 2026-09-04 section of this issue, which records roughly 88.3% adjacent RSS growth against a 10% gate. That section described the state before the repair and was superseded, but sat below newer notes and was mistaken for current. The misleading table has been removed from this issue as part of this edit.

What actually happened: #1094 traced the growth to ShardedCsrIndex::open decoding every shard at open, making resident memory scale with edge count, and closed as completed on 2026-09-07. The 710c6c64 attempt on 2026-09-14 then cleared rss_bounded_or_plateaued at every rung it ran. Memory is not a live blocker. What actually stopped that attempt at S22 was a disk capacity refusal, since resolved by reclamation.

What could make the S26 run fail

Disk, not memory. This is the binding constraint and the only one with a history of refusals. The S26 projected transient peak is roughly two thirds of this host's total root filesystem, before the reserve is subtracted. Confirm admission from a fresh dry run immediately before launch rather than from any recorded figure, and re-confirm after any large build or cache regrows. The margin has moved by tens of gigabytes between samples taken days apart, in both directions.

Wall time. S22 took 1,683 seconds. S26 is sixteen times its edge count. Even perfectly linear scaling lands near 7.5 hours, and the per-rung envelope is four hours. S24 and S25 must run first and will show the real slope, but a time refusal at S24 or S25 is a plausible outcome and should not be read as a regression.

RSS extrapolation is favorable but unproven past S22. Growth across the measured rungs is about 3% and 5% for the single doublings, and 32% across the S20 to S22 four-fold jump. Extrapolating that slope to S26 lands under 1 GB against a 4 GiB per-rung ceiling on a 125 GiB host. The risk is not the magnitude, it is that no rung past 67 million edges has ever been measured, so a new superlinear term would be invisible until it appears.

The run is sequential and stops at the first typed failure. Three rungs must pass in order. A failure at S24 costs the S24 time and yields no S26 evidence. Budget for that rather than treating S26 as a single attempt.

Preconditions to check before launch, none of which are numbers this issue should store. Work root is on the durable ext4 process root, not an overlay. The output directory is fresh. The reserve was recomputed from today's filesystem. Executables are frozen and their digests recorded. The dry run for S24 returns admitted.

Problem

GraphForge still lacks one reproducible Graph500 ladder that proves the ordinary product lifecycle from lower qualification scales through the final billion-edge result. Developer laptops and overlay-rooted VMs are not systems under test. The ladder must expose architectural amplification before an expensive rung and must not compensate by increasing RAM or timeout.

Objective

Execute a sequential qualification and certification ladder on the dedicated Linux bench host OVHC-AGENCY:

S18 -> S19 -> S20 -> S22 -> S24 -> S25 -> S26

Every rung runs the identical ordinary public lifecycle under BenchExec authority with the local-linux-cgroups-v2 profile. Each completed rung produces sanitized evidence; the next rung runs only when the controller admits it from required prior evidence and the explicitly authorized maximum scale permits it.

Designated host (SUT)

  • Hostname: OVHC-AGENCY
  • OS: Ubuntu 26.04 (Linux x86_64)
  • CPUs: 16
  • RAM: 125 GiB
  • Root filesystem: ext4 on NVMe RAID (/dev/md3), approximately 878 GiB total
  • Resource authority: BenchExec cgroups v2 (local-linux-cgroups-v2)
  • Durable --project paths are admitted on this host (process-root ext4); do not run the ladder on overlay-rooted Cloud Agent VMs

Free space is not recorded here. Measure it at launch; see Where the current numbers live.

Requirements

  • Freeze the exact merged commit, generator identity, EF16, seed, host profile, tool versions, and maximum authorized scale.
  • Use the completed public runner/controller (refactor(bench): extract a public-API scale certification runner #955/test(bench): define the complete S18-S26 qualification ladder #956), BenchExec authority (test(bench): enforce resource limits and normalize evidence with BenchExec #957), and ordinary staged-construction path (fix(api): route staged Parquet import through construction sessions #991). Prefer native local admission on this host over disposable Fly provider execution.
  • Keep large datasets, projects, portable exports/imports, and temporary artifacts on this host's local NVMe under a declared work root; do not treat a Mac or overlay VM as the SUT.
  • Generate each scale's deterministic dataset exactly once per ladder attempt and reuse those exact bytes for that rung's ingest, recovery, verification, and any authorized same-commit retry. Do not derive smaller canonical rungs by truncating/remapping a larger-scale dataset.
  • After a rung's evidence is independently accepted, delete that rung's datasets/projects/exports/imports from the host work root before generating the next rung, while retaining only sanitized evidence and immutable identities.
  • Run every rung through generate, ingest, reopen/recount, canonical one-hop/two-hop ordered-LIMIT queries, portable-v2 export, full verify, clean import, imported reopen/recount, and matching post-import queries.
  • Stop at the first typed failure. Never skip a rung, rerun an unchanged failing configuration, or emit a pass for an interrupted rung.
  • Reconcile raw attempts, rejects, duplicates, self-loops, live rows, source/imported counts, construction chunks, application bytes/calls, artifacts, fsyncs, reader calls, and publication work.
  • Record BenchExec process-tree wall/CPU/RSS/physical I/O and GraphForge logical/storage/construction evidence for every phase. Separate process VmHWM/anonymous RSS from file/page-cache.
  • Require bounded or plateauing phase RSS. Continued material RSS growth with edge count is an architectural failure signal.
  • S20 requires conservative adjacent S18/S19 projection within four hours, 4 GiB process RSS, and host disk headroom with a documented reserve.
  • S26 requires completed adjacent S24/S25 projection. Projected transient peak must fit this host's free capacity with a 15% reserve before launch. Peak RSS ≤128 GiB remains only the M5 ceiling, never the default sizing target; prefer measured headroom on this 125 GiB host.
  • Treat the four-hour limit as a per-rung runaway/certification envelope for S18–S22 only, not a universal SLA. S24, S25 and S26 run with no BenchExec time limit so that every phase is measured (maintainer decision 2026-10-05). Their time and RSS projections are recorded but do not refuse admission; storage headroom still does. No rung is stopped on CPU time. Physical I/O is projected and recorded, never an admission gate.
  • Launch each rung only after a sustained quiet-host window (no compiler, benchmark or gf process, low CPU), and record that window in the rung plan.
  • Preserve prior CURRENT on cancellation, corruption, resource failure, or interrupted publication; re-entry must not duplicate rows or work.
  • Validate closed sanitized evidence independently before authorizing the next rung.
  • Tear down and independently verify removal of temporary datasets, projects, exports/imports, and work-root debris after the terminal rung.

Acceptance criteria

  • S18 and S19 complete on OVHC-AGENCY with the full unchanged lifecycle and linear/bounded evidence.
  • Evidence proves one generation per rung, exact dataset identity reuse within the rung, and post-acceptance reclamation before the next scale.
  • The S18/S19 projection admits S20 with documented runtime, RSS, storage, I/O, reader-call, and publication headroom.
  • S20 completes the full source → export → verify → clean import → reopen lifecycle within four hours and 4 GiB process RSS with matching correctness evidence.
  • S22, S24, and S25 run only after their preceding gates pass and preserve bounded/plateauing RSS plus reasonable storage/I/O slopes. (S22 done; S24 and S25 not started.)
  • The S24/S25 projection admits S26 below the host free-capacity transient threshold (15% reserve) with measured RSS/CPU/disk headroom on OVHC-AGENCY.
  • S26 completes with >=1,000,000,000 live persisted edges and matching source/imported counts, queries, fingerprints, and portable-v2 verification.
  • Checked-in evidence contains only versioned sanitized JSON/journals and exact immutable identities; no datasets, graph/package content, credentials, UUID inventories, or absolute host paths beyond agreed sanitized host identity fields.
  • Complete local work-root teardown is proven by independent inventory.
  • Any implementation/evidence/docs PR has exact-head green CI and CI Gate.

Non-goals

Laptop/overlay-VM scale runs, benchmark-only engine paths, weakening queries/durability/verification, increasing RAM/timeouts to conceal amplification, treating bounded memory alone as success, skipping directly to S20/S26, claiming a failed/interrupted rung, or treating this single host as a universal hardware-independent capacity claim.

Recording point-in-time host capacity figures in this issue is also a non-goal. They go stale silently and have already caused one wrong conclusion.

Confirmed benchmark target

Maintainer decision: use Graph500-compliant generated data for GraphForge scale tests, clearly labeled. This release does not claim an official Graph500 BFS/SSSP result or TEPS score. GraphForge's ingest, persistence, query, portable interchange, and recovery measurements remain product lifecycle tests.

Preserve the required Graph500 input semantics and publish the pinned generator identity, SCALE, edge factor, seed, raw tuple counts, and GraphForge stored-edge policy. The Graph500 data standard does not define GraphForge's RSS-growth policy.

Admission uses actual work-root free capacity and the declared reserve. Native run evidence and cleanup satisfy consumers without Fly image/provider requirements or hand-entered throughput certificates. Retain declared resource limits, first-failure stopping, real phase resource evidence, correctness/recovery checks, and one reusable ladder result across the existing completion trackers.

Release inclusion

M5 is part of the approved v0.6.0 scope. This work reaches readiness #1096 and publication #1095 through #900/#745/#735. Keep existing acceptance criteria and reuse the shared OVHC-AGENCY ladder evidence.

M11's broad GDC suites and optional Fly tooling are not prerequisites for this host-native ladder. Preserve required correctness, negative tests, source/imported identity, storage headroom, per-rung evidence, and teardown.

Relationships

S26 start prerequisite: public-operation amplification repairs

Maintainer instruction (2026-10-04): finish #1803 and its eleven native repair children #1421 and #1804–#1813 before launching S26. The epic and all children are in M5; this issue is natively blocked by #1803. Its children cover add_edge rollback snapshots, write initialization, property inventory refresh, transaction backups, relationship MERGE, label-fragment updates, unchanged participant reuse, source registration, derivation validation, inspection helpers and incremental ledger row validation.

#1803 closes on merged repairs and deterministic public-operation correctness/work evidence before S26, so it does not depend on the scale result. Existing adjacent-rung admission, resource/capacity reserves, corruption/refusal and recovery checks, and explicit run authorization remain unchanged. Targeted bounded repair validation is allowed; this dependency is not permission to launch S26. #735 remains the M5 close gate.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions