Skip to content

ci(e2e): classify hosted-runner resource pressure and infrastructure loss #7146

Description

@apurvvkumaria

Parent epic: #7140

Related dependency: #7101 owns general secret-safe phase heartbeats and progress coverage across E2E targets. This issue owns accurate resource attribution, failure classification, and retry policy.

Problem

Current host-memory snapshots can make healthy Linux page cache look like memory exhaustion, while a GitHub-hosted VM can disappear before cleanup, logs, or artifacts identify the cause. Without cgroup, pressure, process, and Docker evidence, maintainers cannot reliably distinguish:

  • application assertion failures;
  • a process or container OOM kill;
  • disk or inode exhaustion;
  • a stalled Docker build; and
  • hosted-runner infrastructure loss.

That uncertainty encourages broad retries, which can hide deterministic regressions.

Scope

  • Extend the secret-safe heartbeat contract with accurate host measurements where available:
    • MemAvailable, Cached, SReclaimable, swap use, and load;
    • cgroup memory.current, peak/limit, memory.events, and OOM/OOM-kill counters;
    • memory and I/O pressure stall information;
    • top process RSS/PSS consumers;
    • Docker container stats, image/build-cache usage, workspace free space, and inode availability; and
    • kernel OOM evidence where the hosted environment permits access.
  • Emit bounded snapshots before and after expensive image/install/rebuild phases and periodically while a phase is active.
  • Produce a machine-readable terminal classification for ordinary failures: assertion, timeout, process OOM, container OOM, disk pressure, or unknown.
  • Define workflow-level signatures for hosted-runner loss when the runner disappears before test cleanup can execute.
  • Permit at most one retry only for confirmed infrastructure-loss signatures. Never retry assertions, deterministic command failures, policy violations, or classified OOM failures.
  • Keep all streamed and uploaded evidence free of command payloads, credentials, tokens, and environment-variable values.

Acceptance criteria

  • Low raw MemFree alone is never classified as OOM.
  • Tests cover classification of application failure, timeout, cgroup OOM, disk exhaustion, runner loss, and ambiguous failure.
  • An ordinary assertion receives zero automatic retries.
  • A confirmed hosted-runner-loss result receives no more than one retry and links the two attempts for diagnosis.
  • The original failure evidence remains available when a retry succeeds.
  • Representative rebuild-hermes and other heavy lanes emit enough evidence to identify the largest host/container memory consumers and Docker disk use.
  • The implementation interoperates with Add secret-safe phase progress diagnostics to every E2E test #7101 rather than creating a second progress-heartbeat framework.

Non-goals

  • Increasing test timeouts.
  • Retrying failures that lack a positive infrastructure-loss classification.
  • Treating larger runners as a substitute for diagnostics.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

area: ciCI workflows, checks, release automation, or GitHub Actionsarea: e2eEnd-to-end tests, nightly failures, or validation infrastructurearea: observabilityLogging, metrics, tracing, diagnostics, or debug outputplatform: containerAffects Docker, containerd, Podman, or imagesv0.0.131Release target

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions