Repository navigation
An eval that completed without a result is an infrastructure failure, not a negative result - #478
Conversation
… bounded uncharged retries, then park
There was a problem hiding this comment.
Round 1 — reviewed head 636863f5 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: 1 blocking, 0 advisory.
1 finding attached to the lines below.
Merged one blocking finding supported by credentials, deployment, general, lifecycle, and prose: exhausted infrastructure retries bypass the wake cap and can launch unbounded uncharged evaluations. Rejected: none; the five opinions describe the same failure mode. Source validation was unavailable because src/outerloop/measure.py is absent from the supplied tree.
# Conflicts: # CHANGELOG.md
There was a problem hiding this comment.
Round 3 — reviewed head 6773ac12 — reviewer hermes/gpt-5.6-terra.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: nothing blocking — 1 advisory note.
1 finding attached to the lines below.
The new retry behavior lacks a candidate COMPLETED-without-result test.
There was a problem hiding this comment.
Round 4 — reviewed head 63d4f4ff — reviewer hermes/gpt-5.6-terra.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: 1 blocking, 0 advisory.
1 finding attached to the lines below.
A crash between retry submission and marker replacement can make infrastructure retries unbounded.
# Conflicts: # CHANGELOG.md
Draft for the maintainer: changes how missing eval results end a run. Fixes #474.
Problem. An eval that finishes without a result raises
EvalError, and the run endsnegative-result. On a live fleet the controller's disk filled. Baseline evals then exited COMPLETED but couldn't writeresult.json, and two runs ended negative with no research outcome.Change (
measure.pyresults()).EvalErrorand messages.infra-retries.jsonbeside the measure'ssubmittedmarker. The count is written atomically and deduplicated by job id, so restarts and repeated deliveries don't reset it. Retries park through the existingcapacity_waitpath, so they aren't charged to the run's GPU budget or retry caps.EvalInfraError(aMeasurementPending) parks the run instead of ending it.outerloop statusshows the job, its state and the missing-result note.After exhaustion (decided in review, maintainer may revisit). An exhausted measure keeps its
submittedmarker, so later wakes and restarts re-read the journal and stay parked without dispatching another eval. The run resumes when an operator fixes the cause and deletes the namedinfra-retries.json. Round 1 of review flagged the earlier "retry once per wake" version as unbounded, because capacity-wait parks wake every 60 seconds. If you'd rather have a timed backoff, it slots in at the same point.Compatibility. The new per-measure file is ignored by older kernels. On rollback, a run parked this way falls back to the old ending on its next wake. Legacy-record and interruption tests are included. The CHANGELOG has an
Upgrading:line.Tests. Baseline and candidate COMPLETED-without-result each re-dispatch and then succeed. A candidate FAILED-without-result keeps the plain
EvalError. After 2 retries the run parks and is not ended. The retry count survives a restart. All 4 mutation checks fail as they should. Full gate: ruff, format, mypy,3216 passed, 12 skipped.Built by Codex (gpt-5.6-terra); cross-reviewed by Claude Opus 5.5, who raised the open question. Advisory review panel to follow.
Slim-down (31c9ac6, 5f14b73). Same behavior, +394 → +181 lines: the 8 new test functions (~40 cases) became two behavior tests;
EvalInfraError(never caught distinctly) removed; the retry journal is self-describing (retries,job_id,note) and written with the existing atomic writer; fixed an off-by-one in the exhaustion count. A failed journal write still parks the run uncharged, and a journaled retry for the same job outranks an unavailable scheduler status. Slim-down built by GLM-5.3 (self-reviewed), reviewed by Codex (two findings, fixed in 5f14b73), final read by Claude Opus 5.5.