Skip to content

CI: windows-latest test job cancelled at the 12-minute ceiling, reported as a failing check #717

Description

@lidge-jun

Area

GitHub Actions CI (.github/workflows/ci.yml).

Summary

The test job's timeout-minutes: 12 is no longer enough headroom for windows-latest, so PR and dev runs are being cancelled mid-Test and surfaced to authors as a failing check. This is a CI budget problem, not a defect in the PRs that trip it.

Evidence from live runs:

Run Branch windows-latest Wall time
30459554635 dev success 11 min
30493348190 dev cancelled —
30497549930 dev cancelled —
30498312557 dev cancelled —
30498333662 dev cancelled —
30498427875 dev cancelled —

The last green Windows run took 11 minutes against a 12-minute ceiling — roughly one minute of margin. Every dev run after it was cancelled.

Job-level confirmation on PR #653 (879efb243, job 90733427766):

conclusion: cancelled
started 2026-07-29T23:11:50Z  completed 2026-07-29T23:23:55Z   (12m05s)
Test: cancelled
GUI tests / Privacy scan / GUI lint / GUI build / CLI help smoke: skipped

The job dies at 12m05s with Test cancelled and every later step skipped, which is the signature of the job timeout rather than a test assertion failure.

Why this is costing review time

gh pr checks renders a cancelled job as fail, so a PR whose code is fine looks broken. Two examples from today:

Reviewers currently have to open gh api .../check-runs and read conclusion on every red Windows job to tell a real failure from a timeout. That is easy to skip, and skipping it either blocks a good PR or hides a real one.

Suggested direction

The immediate unblock is more headroom — the existing code comment already records that 8 minutes was raised to 12 for the same reason, so the ceiling has been chasing suite growth rather than leading it. Worth pairing that with something that stops the ratchet:

  • Raise timeout-minutes for the test job with real margin over the current 11-minute Windows baseline.
  • Look at why Windows is roughly 2.5x slower than Linux (4 min) on the same suite. If a small number of tests dominate, sharding or test.serial on the resource-heavy ones may be cheaper than repeatedly raising the ceiling.
  • Related but separate: the earlier fix(catalog): stop respawning the codex --version probe on every catalog read #610 investigation captured error: EEXIST: file already exists, epoll_ctl plus Cannot call afterEach() after the test run has completed from the Bun runner on Linux CI. I could not find a matching upstream Bun issue, so I am not claiming a known regression — noting it because a runner that crashes or hangs would also present as a timeout.

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, and personal data.

Activity

  1. github-actions commented on Jul 29, 2026

    @github-actions
    Contributor

    Automated translation bookkeeping — detected language: English.

  2. lidge-jun commented on Jul 29, 2026

    @lidge-jun
    OwnerAuthor

    Correcting one part of my own report, and sharpening the diagnosis with per-job timings.

    What I got wrong: I listed five cancelled dev runs as evidence of the timeout. They are not. Their total wall times were 0.6, 0.8, and 1.9 minutes — far too short to be a 12-minute timeout. Those were cancel-in-progress: true from the concurrency group firing as three PRs merged in quick succession. Normal behaviour, not this bug.

    What holds, and is actually the sharper finding: the real timeout is reproducible and the margin is far thinner than "roughly one minute".

    Run PR windows-latest Job wall time
    30480687886 #711 rerun success 11.8 min
    30476667108 #653 rerun cancelled 12.0 min
    30459554635 dev success 11.8 min

    Against timeout-minutes: 12, a green Windows job finishes at 11.8 minutes. That is roughly 12 seconds of headroom. #653 crossed it and was killed at the Test step with every later step skipped; #711 landed just under and passed. Same workflow, same week, and the outcome is decided by runner variance rather than by the code under review.

    For contrast, the other two platforms on that same green run:

    ubuntu-latest    4.6 min
    macos-latest     5.6 min
    windows-latest  11.8 min
    

    So Windows is ~2.5x Linux on an identical suite, and it is the only platform anywhere near the ceiling.

    This makes the check untrustworthy in both directions, which is the part that costs review time. A red Windows job can mean a real failure (#653's first run was a genuine conclusion: failure), a concurrency cancel, or a timeout — and gh pr checks renders all three as fail. Telling them apart currently requires gh api repos/lidge-jun/opencodex/commits/<sha>/check-runs and reading conclusion, plus job timings to separate a timeout from a concurrency cancel.

    Raising the ceiling fixes the immediate flakiness. It does not stop the ratchet — the code comment already records 8 → 12 for this same reason, and 12 is now spent. The durable question is why Windows needs 2.5x, and whether a few resource-heavy suites dominate that gap.

  3. lidge-jun commented on Jul 30, 2026

    @lidge-jun
    OwnerAuthor

    Measured where the Windows time actually goes, and tested three ways to cut it. Posting the numbers because two of the three obvious options are dead ends, and it is worth not re-discovering that.

    The whole problem is one step

    Per-step timings from the green run 30459554635 (windows-latest):

    Test                       581s   ← 82% of the job
    Install dependencies        37s
    GUI tests                   31s
    GUI lint                    21s
    Typecheck                   11s
    Checkout                     9s
    GUI build                    9s
    Privacy scan                 4s
    everything else            < 5s
    

    Total 708s. Test is 581s; every other step combined is under two minutes. Same run on ubuntu-latest: Test 206s. So the 2.5x platform gap is entirely inside bun test, and nothing else in the workflow is worth touching.

    Local baseline for reference: bun test --isolate tests/ runs 5990 tests across 430 files in 220.87s at 76% CPU — a single process, with the rest of the cores idle. That idle capacity is the opportunity.

    Option 1: drop --isolate — rejected

    Not a waste to remove; it is load-bearing. Without it the suite booted the proxy server 307 times into a shared global and hung past 11 minutes before I killed it. --isolate is what keeps leaked handles from one file out of the next.

    Option 2: --parallel=N — rejected as-is

    bun test --parallel=4 produced 7 failures, all in provider-management and credential-separation tests. Those same tests pass cleanly on their own (tests/management-provider-validation.test.ts → 35 pass, 0 fail). So it is contention, not defects.

    Option 3: --shard=N/M — viable, but blocked by the same root cause

    bun test --isolate --shard=1/3 tests/ ran 144 files in 89s, about 40% of the full wall time. Splitting into three CI jobs would put Windows near 3-4 minutes instead of 10-12. But the same provider-management failures appear, for the same reason.

    What actually blocks parallelism

    27 test files pin a fixed temp directory and point OPENCODEX_HOME at it:

    // tests/management-provider-validation.test.ts:44
    const TEST_DIR = join(import.meta.dir, ".tmp-server-auth-test");
    ...
    process.env.OPENCODEX_HOME = TEST_DIR;

    One directory, one name, shared by every worker that touches it. Sequential runs never notice; two workers do.

    The fix pattern already exists in this repo — tests/helpers/isolated-codex-home.ts uses mkdtempSync(join(tmpdir(), prefix)) and restores the previous env on teardown. It currently covers CODEX_HOME only, so the equivalent for OPENCODEX_HOME would need adding, but the shape is established and 125 files already use per-run temp dirs.

    Suggested order

    1. Give those 27 files a per-run temp dir (mirror the existing helper, extended to OPENCODEX_HOME). This is the precondition for everything else and is worth doing on its own — a fixed shared path under tests/ is a latent flake even today.
    2. Then shard the test job 3 ways. Expected Windows wall time 3-4 min, which also retires the timeout pressure this issue was opened about.
    3. Optionally revisit --parallel after step 1; sharding across jobs is the safer first move since each job stays single-process.

    One thing that will not help: trimming the slow tests. 55 tests take over a second and account for 112.9s of 207.5s, but the slowest ones are deliberate stall/timeout assertions — honors slow_down without failing the device flow (7.0s), stalled 400 body timeout never authorizes a pool retry (5.1s), stalled passthrough JSON is canceled at five seconds (5.0s). They have to really wait. Their cost is a reason to run them in parallel with other work, not to shorten them.

  4. lidge-jun commented on Jul 30, 2026

    @lidge-jun
    OwnerAuthor

    Resolved. .github/workflows/ci.yml:61 now sets timeout-minutes: 20 for the test job (raised in 5bf66df26, with a76bf413a following up), so the ceiling is no longer the binding constraint.

    Confirmed against live runs from this session rather than from the config alone. The longest windows-latest job I watched ran 12 minutes (01:50:42Z → 02:02:41Z) and completed success — it would have been cancelled under the old 12-minute ceiling with roughly zero margin. Green dev run 30505367421 also has all six jobs passing, including windows-latest.

    Two adjacent findings from chasing this, so the next person does not misread them as the same problem:

    The old cancellations were genuinely a budget issue, but not every red windows-latest was. While landing #646 I hit three distinct causes in a row — this timeout ceiling, then the kiro platform regression (#718), then a Bun runtime panic(thread …): Internal assertion failure that cleared on rerun. Worth checking the actual job conclusion and the failing step before assuming the ceiling.

    Related trap: gh pr checks reports fail for a cancelled job, so a cancelled Windows run looks identical to a real test failure from the PR page. gh api repos/.../actions/jobs/<id> gives the true conclusion.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions