Skip to content

test(bench): define the complete S18-S26 qualification ladder #956

Description

@DecisionNerd

Objective

Define the complete Graph500 qualification ladder as declarative ReFrame/BenchExec profiles and a fail-closed controller, without executing provider scale runs from this issue.

Ladder

S18 -> S19 -> S20 -> S22 -> S24 -> S25 -> S26

  • S18/S19 are lower qualification rungs, not developer-laptop benchmarks.
  • S20 is projected only from adjacent completed S18/S19 evidence.
  • S22 is separately gated by completed lower evidence.
  • S26 is projected only from adjacent completed S24/S25 evidence.
  • Every rung uses the same ordinary source -> ingest -> reopen/recount -> canonical queries -> export -> verify -> clean import -> reopen proof lifecycle.

Requirements

  • Add distinct S18, S19, S20, S22, S24, S25, and S26 profiles.
  • Pin generator identity, EF16, seed, public lifecycle phases, evidence schemas, limits, and headroom policy without provider resource IDs.
  • Select only the first unexecuted rung; a failed or incomplete rung authorizes nothing larger.
  • Project wall time, phase/process-tree RSS, retained/transient storage, logical/physical I/O, reader calls, and publication work from the latest required adjacent completed rungs.
  • Treat sustained material RSS growth with edge count as an architectural refusal signal. Persistent size is expected to be disk-bound.
  • Require measured provider throughput/capacity for physical I/O, reader calls, and publication work; missing metrics fail closed.
  • Stage immutable executable copies and verify exact SHA-256 identities immediately before execution.
  • Emit only closed, sanitized, schema-valid evidence and typed first-failure results.
  • Keep profiles outside normal product CI. Tiny fixtures and dry-runs may run locally.
  • The disk-constrained macOS controller host must not build or retain large Linux targets, Graph500 datasets, project copies, OCI layers, or scale evidence payloads.

Acceptance criteria

  • Closed schemas and tiny fixtures validate all seven profiles, phase order, executable identities, selection, refusal, typed failures, and evidence sanitization.
  • S24 and S25 are explicit executable profiles; S26 cannot be admitted without both.
  • Conservative projection independently checks the four-hour rung limit, S20 4 GiB envelope, Fly 500 GB maximum with 15% storage reserve, measured I/O/reader/publication capacity, and bounded/plateauing RSS.
  • The controller stages a selected profile and immutable executables, invokes BenchExec only after native-Linux admission, validates normalized output, and never provisions Fly itself.
  • Missing ordinary-path metric receipts produce a typed refusal rather than inferred counters.
  • No scale rung runs in normal CI or as issue-close evidence for this construction issue.

BDD completion scenario

Given a clean checkout, the complete seven-rung profile set, and fixture evidence
When the controller selects or projects the next rung
Then it chooses only the first authorized rung
And refuses absent metrics, non-adjacent projections, identity drift, resource-envelope violations, or prior failure.

Non-goals

Executing the real Fly ladder, provisioning provider resources, weakening canonical queries/durability/verification, benchmark-only engine hooks, laptop scale measurements, or claiming official Graph500 submission results.

Relationships

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    testingTest coverage and testing infrastructure

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions