Skip to content

An eval that completed without a result is an infrastructure failure, not a negative result - #478

Merged
renmengye merged 7 commits into
mainfrom
fix/no-result-is-infra
Oct 11, 2026
Merged

renmengye merged 7 commits into
mainfrom
fix/no-result-is-infra

Conversation

@renmengye

@renmengye renmengye commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Draft for the maintainer: changes how missing eval results end a run. Fixes #474.

Problem. An eval that finishes without a result raises EvalError, and the run ends negative-result. On a live fleet the controller's disk filled. Baseline evals then exited COMPLETED but couldn't write result.json, and two runs ended negative with no research outcome.

Change (measure.py results()).

  • Classification. A missing result counts as an infrastructure failure when the measure is a baseline, including suite baselines (they never run agent code), or when the job ended COMPLETED / exit 0. A crashing candidate exits non-zero. FAILED, OOM, TIMEOUT and signal deaths of candidates keep today's EvalError and messages.
  • Retries. Up to 2 re-dispatches, counted in infra-retries.json beside the measure's submitted marker. The count is written atomically and deduplicated by job id, so restarts and repeated deliveries don't reset it. Retries park through the existing capacity_wait path, so they aren't charged to the run's GPU budget or retry caps.
  • Exhaustion. EvalInfraError (a MeasurementPending) parks the run instead of ending it. outerloop status shows the job, its state and the missing-result note.

After exhaustion (decided in review, maintainer may revisit). An exhausted measure keeps its submitted marker, so later wakes and restarts re-read the journal and stay parked without dispatching another eval. The run resumes when an operator fixes the cause and deletes the named infra-retries.json. Round 1 of review flagged the earlier "retry once per wake" version as unbounded, because capacity-wait parks wake every 60 seconds. If you'd rather have a timed backoff, it slots in at the same point.

Compatibility. The new per-measure file is ignored by older kernels. On rollback, a run parked this way falls back to the old ending on its next wake. Legacy-record and interruption tests are included. The CHANGELOG has an Upgrading: line.

Tests. Baseline and candidate COMPLETED-without-result each re-dispatch and then succeed. A candidate FAILED-without-result keeps the plain EvalError. After 2 retries the run parks and is not ended. The retry count survives a restart. All 4 mutation checks fail as they should. Full gate: ruff, format, mypy, 3216 passed, 12 skipped.

Built by Codex (gpt-5.6-terra); cross-reviewed by Claude Opus 5.5, who raised the open question. Advisory review panel to follow.

Slim-down (31c9ac6, 5f14b73). Same behavior, +394 → +181 lines: the 8 new test functions (~40 cases) became two behavior tests; EvalInfraError (never caught distinctly) removed; the retry journal is self-describing (retries, job_id, note) and written with the existing atomic writer; fixed an off-by-one in the exhaustion count. A failed journal write still parks the run uncharged, and a journaled retry for the same job outranks an unavailable scheduler status. Slim-down built by GLM-5.3 (self-reviewed), reviewed by Codex (two findings, fixed in 5f14b73), final read by Claude Opus 5.5.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 636863f5 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: 1 blocking, 0 advisory.

1 finding attached to the lines below.

Merged one blocking finding supported by credentials, deployment, general, lifecycle, and prose: exhausted infrastructure retries bypass the wake cap and can launch unbounded uncharged evaluations. Rejected: none; the five opinions describe the same failure mode. Source validation was unavailable because src/outerloop/measure.py is absent from the supplied tree.

Comment thread src/outerloop/measure.py
@renmengye renmengye added the outerloop:review re-request the advisory review label Oct 4, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 4d489603 — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: no defects found.

@renmengye renmengye added outerloop:review re-request the advisory review and removed outerloop:review re-request the advisory review labels Oct 11, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 — reviewed head 5f14b736 — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: no defects found.

@renmengye renmengye added outerloop:review re-request the advisory review and removed outerloop:review re-request the advisory review labels Oct 11, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 3 — reviewed head 6773ac12 — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: nothing blocking — 1 advisory note.

1 finding attached to the lines below.

The new retry behavior lacks a candidate COMPLETED-without-result test.

Comment thread tests/test_measure.py Outdated
@renmengye renmengye added outerloop:review re-request the advisory review and removed outerloop:review re-request the advisory review labels Oct 11, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 4 — reviewed head 63d4f4ff — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: 1 blocking, 0 advisory.

1 finding attached to the lines below.

A crash between retry submission and marker replacement can make infrastructure retries unbounded.

Comment thread src/outerloop/measure.py
@renmengye
renmengye marked this pull request as ready for review October 11, 2026 13:46
@renmengye
renmengye merged commit 0bd5338 into main Oct 11, 2026
1 check passed
@renmengye
renmengye deleted the fix/no-result-is-infra branch October 11, 2026 13:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

outerloop:review re-request the advisory review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An eval that finishes without a result ends the run as negative-result

1 participant