Repository navigation
CI: windows-latest test job cancelled at the 12-minute ceiling, reported as a failing check #717
Description
Activity
github-actions commented
on Jul 29, 2026 on Jul 29, 2026 – with GitHub ActionsContributorMore actionsAutomated translation bookkeeping — detected language: English.
Correcting one part of my own report, and sharpening the diagnosis with per-job timings.
What I got wrong: I listed five cancelled
devruns as evidence of the timeout. They are not. Their total wall times were 0.6, 0.8, and 1.9 minutes — far too short to be a 12-minute timeout. Those werecancel-in-progress: truefrom the concurrency group firing as three PRs merged in quick succession. Normal behaviour, not this bug.What holds, and is actually the sharper finding: the real timeout is reproducible and the margin is far thinner than "roughly one minute".
Run PR windows-latest Job wall time 30480687886 #711 rerun success 11.8 min 30476667108 #653 rerun cancelled 12.0 min 30459554635 devsuccess 11.8 min Against
timeout-minutes: 12, a green Windows job finishes at 11.8 minutes. That is roughly 12 seconds of headroom. #653 crossed it and was killed at theTeststep with every later step skipped; #711 landed just under and passed. Same workflow, same week, and the outcome is decided by runner variance rather than by the code under review.For contrast, the other two platforms on that same green run:
ubuntu-latest 4.6 min macos-latest 5.6 min windows-latest 11.8 minSo Windows is ~2.5x Linux on an identical suite, and it is the only platform anywhere near the ceiling.
This makes the check untrustworthy in both directions, which is the part that costs review time. A red Windows job can mean a real failure (#653's first run was a genuine
conclusion: failure), a concurrency cancel, or a timeout — andgh pr checksrenders all three asfail. Telling them apart currently requiresgh api repos/lidge-jun/opencodex/commits/<sha>/check-runsand readingconclusion, plus job timings to separate a timeout from a concurrency cancel.Raising the ceiling fixes the immediate flakiness. It does not stop the ratchet — the code comment already records 8 → 12 for this same reason, and 12 is now spent. The durable question is why Windows needs 2.5x, and whether a few resource-heavy suites dominate that gap.
Measured where the Windows time actually goes, and tested three ways to cut it. Posting the numbers because two of the three obvious options are dead ends, and it is worth not re-discovering that.
The whole problem is one step
Per-step timings from the green run 30459554635 (
windows-latest):Test 581s ← 82% of the job Install dependencies 37s GUI tests 31s GUI lint 21s Typecheck 11s Checkout 9s GUI build 9s Privacy scan 4s everything else < 5sTotal 708s.
Testis 581s; every other step combined is under two minutes. Same run onubuntu-latest:Test206s. So the 2.5x platform gap is entirely insidebun test, and nothing else in the workflow is worth touching.Local baseline for reference:
bun test --isolate tests/runs 5990 tests across 430 files in 220.87s at 76% CPU — a single process, with the rest of the cores idle. That idle capacity is the opportunity.Option 1: drop
--isolate— rejectedNot a waste to remove; it is load-bearing. Without it the suite booted the proxy server 307 times into a shared global and hung past 11 minutes before I killed it.
--isolateis what keeps leaked handles from one file out of the next.Option 2:
--parallel=N— rejected as-isbun test --parallel=4produced 7 failures, all in provider-management and credential-separation tests. Those same tests pass cleanly on their own (tests/management-provider-validation.test.ts→ 35 pass, 0 fail). So it is contention, not defects.Option 3:
--shard=N/M— viable, but blocked by the same root causebun test --isolate --shard=1/3 tests/ran 144 files in 89s, about 40% of the full wall time. Splitting into three CI jobs would put Windows near 3-4 minutes instead of 10-12. But the same provider-management failures appear, for the same reason.What actually blocks parallelism
27 test files pin a fixed temp directory and point
OPENCODEX_HOMEat it:// tests/management-provider-validation.test.ts:44 const TEST_DIR = join(import.meta.dir, ".tmp-server-auth-test"); ... process.env.OPENCODEX_HOME = TEST_DIR;
One directory, one name, shared by every worker that touches it. Sequential runs never notice; two workers do.
The fix pattern already exists in this repo —
tests/helpers/isolated-codex-home.tsusesmkdtempSync(join(tmpdir(), prefix))and restores the previous env on teardown. It currently coversCODEX_HOMEonly, so the equivalent forOPENCODEX_HOMEwould need adding, but the shape is established and 125 files already use per-run temp dirs.Suggested order
- Give those 27 files a per-run temp dir (mirror the existing helper, extended to
OPENCODEX_HOME). This is the precondition for everything else and is worth doing on its own — a fixed shared path undertests/is a latent flake even today. - Then shard the
testjob 3 ways. Expected Windows wall time 3-4 min, which also retires the timeout pressure this issue was opened about. - Optionally revisit
--parallelafter step 1; sharding across jobs is the safer first move since each job stays single-process.
One thing that will not help: trimming the slow tests. 55 tests take over a second and account for 112.9s of 207.5s, but the slowest ones are deliberate stall/timeout assertions —
honors slow_down without failing the device flow(7.0s),stalled 400 body timeout never authorizes a pool retry(5.1s),stalled passthrough JSON is canceled at five seconds(5.0s). They have to really wait. Their cost is a reason to run them in parallel with other work, not to shorten them.- Give those 27 files a per-run temp dir (mirror the existing helper, extended to
- added a commit that references this issue
on Jul 30, 2026 Resolved.
.github/workflows/ci.yml:61now setstimeout-minutes: 20for the test job (raised in5bf66df26, witha76bf413afollowing up), so the ceiling is no longer the binding constraint.Confirmed against live runs from this session rather than from the config alone. The longest
windows-latestjob I watched ran 12 minutes (01:50:42Z→02:02:41Z) and completedsuccess— it would have been cancelled under the old 12-minute ceiling with roughly zero margin. Greendevrun 30505367421 also has all six jobs passing, includingwindows-latest.Two adjacent findings from chasing this, so the next person does not misread them as the same problem:
The old cancellations were genuinely a budget issue, but not every red
windows-latestwas. While landing #646 I hit three distinct causes in a row — this timeout ceiling, then the kiro platform regression (#718), then a Bun runtimepanic(thread …): Internal assertion failurethat cleared on rerun. Worth checking the actual jobconclusionand the failing step before assuming the ceiling.Related trap:
gh pr checksreportsfailfor a cancelled job, so a cancelled Windows run looks identical to a real test failure from the PR page.gh api repos/.../actions/jobs/<id>gives the trueconclusion.- added a commit that references this issue
on Sep 17, 2026
Area
GitHub Actions CI (
.github/workflows/ci.yml).Summary
The
testjob'stimeout-minutes: 12is no longer enough headroom forwindows-latest, so PR anddevruns are being cancelled mid-Testand surfaced to authors as a failing check. This is a CI budget problem, not a defect in the PRs that trip it.Evidence from live runs:
devdevdevdevdevdevThe last green Windows run took 11 minutes against a 12-minute ceiling — roughly one minute of margin. Every
devrun after it was cancelled.Job-level confirmation on PR #653 (
879efb243, job90733427766):The job dies at 12m05s with
Testcancelled and every later step skipped, which is the signature of the job timeout rather than a test assertion failure.Why this is costing review time
gh pr checksrenders a cancelled job asfail, so a PR whose code is fine looks broken. Two examples from today:windows-latest fail. Theconclusionwascancelled; a rerun passed and it merged cleanly (d24c5233f).failure, but the rerun came backcancelledat 12m05s. Its focused suites pass locally (38 pass) withbun x tsc --noEmitclean, so the preset itself is not implicated.Reviewers currently have to open
gh api .../check-runsand readconclusionon every red Windows job to tell a real failure from a timeout. That is easy to skip, and skipping it either blocks a good PR or hides a real one.Suggested direction
The immediate unblock is more headroom — the existing code comment already records that 8 minutes was raised to 12 for the same reason, so the ceiling has been chasing suite growth rather than leading it. Worth pairing that with something that stops the ratchet:
timeout-minutesfor thetestjob with real margin over the current 11-minute Windows baseline.test.serialon the resource-heavy ones may be cheaper than repeatedly raising the ceiling.error: EEXIST: file already exists, epoll_ctlplusCannot call afterEach() after the test run has completedfrom the Bun runner on Linux CI. I could not find a matching upstream Bun issue, so I am not claiming a known regression — noting it because a runner that crashes or hangs would also present as a timeout.Checks