Repository navigation
ci(ios): the iOS simulator smoke lane is failing on main across several rotating signatures #2491
Description
Activity
Keep the failures separated by signature. #2493 addresses the WebView landmark probe and the wait that gives up on RUNNER_BUSY; it does not resolve the cold-capture, scrolling or build-timeout cases. Different failures on the same commit justify investigation, but do not establish that every red run is unrelated to a PR. Confirm each current signature against main and keep this issue open for the remaining cases after #2493 lands.
Root cause for row 1 (
wait for the WebView page to expose its link)Traced it end to end on the newest occurrence, job
103282380298(PR #2448 head28c6e86bee). It is a lane defect, not a change-under-test defect: the failure is the same signature as main's34592935964, the log has zero system-surface mentions, and the mechanism below is platform-wide and predates the branch.The failing step never polled. The error JSON carries none of the wait evidence fields (
reason,timeoutMs,polls,captures,waitedMs,readableCaptures). It is the raw capture error:command: agent-device wait text Jump to form 20000 ... "code": "RUNNER_BUSY", "message": "The iOS runner is still finishing a previous command that exceeded its execution watchdog (usually an accessibility capture on a heavy or animating screen)." "hint": "Wait a few seconds and retry. ..." details: { command: "snapshot", lifecycleState: "failed", recovery: "runner_reported_failure" }So: a WebView accessibility capture exceeded the runner's execution watchdog, the next capture got
RUNNER_BUSY, andwaitgave up on its first poll with 20 s of budget unspent.Why it gives up, in three steps.
-
The runner error is built as retriable.
packages/platform-apple/src/runner/runner-session.ts:920setsretriable: trueforRUNNER_BUSY, and the code is inDIAGNOSTIC_ONLY_RUNNER_ERROR_CODESso it staysCOMMAND_FAILEDwithdetails.runnerErrorCodepreserved. -
Nothing retries on that.
retriableis a wire classification only:request-router.tsandrequest-finalization.tscopy it outward to the client, and no host-side caller consults it. The only code that readsdetails.runnerErrorCodefor a policy decision ispackages/platform-apple/src/alert.ts:149, and it readsALERT_NOT_FOUND.RUNNER_BUSYappears in that file and inalert-contract.tsonly as prose precedent for the pattern. -
The wait poll loop has exactly one error-tolerance channel, and it is Android-only.
wait textbuilds its polling atsrc/commands/interaction/runtime/wait-text.ts:19with no classification argument, socreateWaitPollingdefaults toisUnreadableCaptureContentError, which inspectsdetails.androidSnapshotHelperFailureReasonand nothing else. Any other throw propagates out ofunreadable.attemptand out of the whole wait.
The
waitfamily therefore has a budget it cannot spend on the one error whose own hint says to wait and retry. The lane is a heavy animating WebView, which is precisely the screen that trips the watchdog, so this row will keep rotating back until the tolerance channel admits a retriable runner stall.Filed as #2496 for the fix. Recording it here so row 1 has a cause rather than a count.
-
Correction to my comment above, and the fix is already in flight: #2493.
Two things in my diagnosis were wrong or incomplete.
The trigger is not "a heavy animating WebView trips the watchdog." It is a landmark mismatch.
acceptDeepLinkConfirmationIfPresenthad a hard-coded Automation-lab readiness landmark, so on any other route it could never match and always fell through to itsalert getprobe. An XCTest alert query against a liveWKWebViewexceeds the runner's 30 s main-thread execution watchdog, and that is what leaves the runner refusing later commands asRUNNER_BUSY.smoke:regular-visible-depth-frontiercarried the same mismatch and the same doomed probe. So row 1 is a scenario-authoring defect first, and the wait's surrender is what turned it into an opaque failure.There is a second classification path I did not find. A runner error recovered from the lifecycle journal after a lost transport response was built with a bare
toAppErrorCode, so on that pathRUNNER_BUSYreached callers as a wire code with noretriableflag at all.#2493 fixes all three: the route landmark, the wait ride-out, and one
classifyRunnerReportedErrorbehind both paths. It carries live before/after validation on a booted simulator. My wait-poll trace above is accurate as a mechanism but is superseded as a root cause. I also filed #2496 before finding #2493 and have closed it as a duplicate.Row 1 should be considered owned by #2493 rather than open here.
Keep row 1 assigned to #2493, but not resolved yet: its latest exact-head iOS smoke still fails at the WebView page wait. The error now has the corrected COMMAND_FAILED / retriable classification, which confirms that part of the fix reached the lane; successful observation of the page is still missing. The other signatures remain separate work.
- added 7 commits that reference this issue
on Sep 15, 2026 Today's signature on the iOS smoke lane, hitting every PR I have open plus two unrelated ones:
AssertionError: regular depth-1 snapshot must disclose the Simulator AX bridge evidence gap at assertSimulatorBridgeSnapshot (test/integration/ios-simulator-e2e/live-snapshot-depth-frontier.ts:126)Runs: 35862236341 (#2804), 35862934402 (#2811), 35862504224 (#2814), 35862792165 (
fix/scroll-movement-disclosure), 35845257256 (fix/native-stack-bridge-hittability). The three PRs of mine touch the daemon resend policy, the daemon-client takeover decision, and two runner log lines respectively; none change the snapshot backend plan.In the failing runs the regular depth-1 snapshot came back from the XCTest tree (
snapshotDiagnostics.backends: {xctest: 54}) with the slow-snapshot warning (p95 3426ms, max 5626ms), so the bridge tier is not the one answering under that load, and the assertion that expects bridge evidence fails. Looks like host load on the runners rather than a code regression; the same job passes on rerun (#2804 had one pass and one fail on the same head).Status after a census of 123 failed iOS Smoke jobs (09-14 to 09-23; main red on 17 of 59 runs). Each rotating signature now has a named mechanism and a merged fix:
Signature Failures Fix must disclose the Simulator AX bridge evidence gap22 (all on 09-23) #2832: #2775 had merged with a stale E2E assertion first wait for Agent Device Testerafteropen --relaunch(wait_capture_stalled)15 #2838: openjoins the in-flight discoverywait for Automation labafter the deep-link "Open in…" prompt12 #2852: runner reads never launch a stopped session app automation-longpress did not become visible20 #2839: reads after an unsettled gesture carry unsettledGestureand the helper re-readsalert-observation XCTests 13 #2843 (closes #2546) testAbandonedTreeCapture…8 #2847 Still open or unaddressed: the webview-link wait (1 sighting in the window), cold
app.launch()timeouts in the XCTest lane (4), #2343 (wait startup accounting), and the follow-ups #2853, #2856 and #2862. I'm keeping this open until main shows a run of consecutive green iOS Smoke results. That is also when iOS Smoke could become a required check, since #2775 merged red.- added 8 commits that reference this issue
on Sep 24, 2026 Fresh attribution pass (2026-10-08, commits
a513106f7…24b2ce638)I pulled the failing-step payloads for every recent
ios.ymlmain failure (gh run view <id> --repo callstack/agent-device --json jobs, thengh run download <id> --repo callstack/agent-device -n "ios-e2e-simulator-summary-*"/-n "ios-e2e-simulator-replay-*"and readwork/ios-steps/<scenario>/failed-step.txt). All 10 recent main failures now resolve to known signatures plus three new ones:Run Signature Attribution 37459193672, 37145696803 snapshotQualitywarning onsnapshot --base64#3328 — misasserted test (product always stamps a strategy on the circuit-disabled XCTest fallback; the AX bridge publishes no verdict). Fix: #3336 37419258095, 37187537985 Settings replay: wait text Automation labdeadline exceeded,@preset-detailnot found, attempt 3/3, step 4/11The #2948-adjacent Settings replay wait — recorded here so it is no longer unattributed (owned under the #2948 follow-up work) 37356199982 ( wait_runner_restart_exhausted), 36454269522 (wait_capture_stalled), 36397945147 + 36247999237 (wait_readiness_exhausted), 36298519121 (wait_readiness_exhausted, runner restart mid-run)hard failure on wait text "Automation lab"insmoke:automation-inputE2E stepInfra-flaked transport/observation waits — the class the retry policy in #3336 now re-issues once 36115336398 gesture-pan replay TIMEOUT+timeout_cleanup_pending(runner 310s, wall 452s)New, untracked — replay-timeout class, cleanup pending after command timeout 36127405340 xcrun --sdk iphoneos --show-sdk-versionETIMEDOUTduring preflightTracked: #2422 37660079219 alert dismiss→MAIN_THREAD_TIMEOUT(phaseafter_command_dispatch) onsmoke:interaction-registrystep 10/17New, untracked — #2782 (finished-at-boundary misreport) is closed and does not cover this; looks like a genuine main-thread stall during the interaction-registry scenario; adjacent to #2956/#2546 alert flakiness 36916418541, 36472317113 run cancelled mid-E2E with wait_runner_restart_exhausted(UI-interaction scenario) atretriable: truewhile the job was observation-preventedNot a scenario verdict — cancelled runs whose last step carried a retriable wait reason; same transport/observation class as the row above Two earlier entries in the sample (37043315673, 37033991194) turned out to be
pull_request-event runs, not push-event main runs — excluded.Measured failure rate
Sample: last 200
ios.ymlruns onmain, window 2026-09-25 → 2026-10-08.gh run list --repo callstack/agent-device --workflow ios.yml --branch main --limit 200 \ --json databaseId,conclusion,event,createdAt gh run list --repo callstack/agent-device --workflow "iOS Backward Compatibility" --branch main --limit 100 \ --json databaseId,conclusion,event,createdAt- 191 push-event
ios.ymlmain runs: 16 failures / 64 successes / 111 cancelled → failure rate among completed push runs ≈ 20% (16/80). - Monthly split: September 9 fails / 32 completed ≈ 28%; October 7 / 48 ≈ 15%.
- Backward-compat lane: all non-failing (32 successes, 66 cancelled).
The cancellations are dominated by concurrency-group supersedes (new pushes), so the meaningful denominator is completed runs only.
Retry policy decision (#3336)
The lane's two replay steps already run under
--retries 2(node test runner); the E2E steps had no layer, so every transport/observation miss ruled red immediately. #3336 adds a shared, opt-in harness seam (reattemptInfrastructureMissinlive-device-e2e/runtime.ts) that re-issues one failed step, enabled by the iOS lane only when the miss is typed as infrastructure:error.retriable === trueAND wait reason ∈ {wait_capture_stalled,wait_runner_restart_exhausted,wait_readiness_exhausted}. No error-text matching.Non-retriable (stays red on first failure):
wait_target_absent,wait_deadline_exceeded, wrong asserted values, runner crashes, and any non-wait failure;allowFailure/expectFailuresteps are exempt. Retried misses still print their first failure before the re-issue; the lane summary reports the count. Bounded at one re-issue per step,INFRASTRUCTURE_MISS_ISSUES = 2in the runtime.Known limitation: the re-issue shares the same runner/session, so a runner-dead miss can re-fail unchanged; per-scenario fresh-runner teardown is the lane-rewrite follow-up this feeds into.
- 191 push-event
Correction to the table above (caught re-checking the downloaded payloads before filing #3337):
- 37660079219: the
MAIN_THREAD_TIMEOUTonalert dismissoccurred insmoke:automation-input(step 55, "dismiss native alert"), notsmoke:interaction-registry. Typed details:reason: runner_main_thread_timeout,dispatched: "unknown",runnerFailureReason: runner_main_thread_execution_timeout. The failing run's step history contains onlybootstrap,smoke:inventory-install, andsmoke:automation-input. - 36115336398: the failing step was "Run gesture pan-duration smoke replay" —
examples/test-app/replays/gesture-pan-duration.ad, attempt 1/3,TIMEOUT after 180000mson awaitcommand withtimeoutCleanupPending: true, junit wall time 182s, and a subsequent** BUILD INTERRUPTED **in that attempt's runner log. The settings replay in the same run passed (7 steps replayed). The "runner 310s / wall 452s" and "12 pass / 6 fail" figures quoted in the original row do not exist in this run's artifacts — disregard them.
Everything else in the comment stands; both corrections are reflected in #3337.
- 37660079219: the
Ledger update: a fourth untracked signature appeared on a PR lane (not main) —
smoke:webview-remote-contentcold-start overrun,wait_deadline_exceededonwait text "Jump to form" 20000with the page content present in the post-failure snapshot (run 37843633972, PR #3336 lane). Filed as #3343. Also filed: #3342 for the preflightdaemon_startup_failedPR-lane class (two branches, identical signature; did not reproduce on re-run at 98321a2). The #3336 lane-policy intentionally leaves both red until their owners land.Ledger correction to my previous comment: the webview cold-start signature (
wait_deadline_exceededwith high readableCaptures, content present in the post-failure snapshot) is now recorded in #3337 as its third scenario-level signature (per the #3336 coordinator pass); my standalone #3343 is closed as a duplicate with a forwarding pointer. The preflightdaemon_startup_failedclass stays in #3342, now corroborated by a third cross-branch instance (#3331 run 37833483850 attempt 1, 19:44Z) and non-reproduction on head 98321a2.Attribution update from the PR-supervision coordinator: this signature now also reaches the
Preflight iOS runner through public CLIstep, where it fails the job before any test runs.Instance (2026-10-09T00:40:20Z, attempt 1): run
37865442839, job113611128846, headce40232e3onfix/android-shutdown-ime-flush-window(PR #3331, an Android-only diff — it cannot touch an Apple path). Failed step:Preflight iOS runner through public CLI. Payload:"code": "COMMAND_FAILED" "message": "xcrun timed out after 15000ms" "hint": "Retry with --debug and inspect diagnostics log for details." "diagnosticId": "mv08mt9n-400b51b8"Same signature, new location. This table already records
xcrun --sdk iphonesimulator --show-sdk-version (spawnSync xcrun ETIMEDOUT)after** TEST BUILD SUCCEEDED **in the runner-build step, matching #2422. This time the identical probe call is what fails, but it runs earlier: the preflight step is the firstprepare ios-runneron the lane, so the failure aborts the job with zero smoke steps executed and nostep-history.json, which is why it reads as a different class than it is.Call site, for whoever owns the budget:
exec.ts:368builds thetimed out after <n>msmessage, and the only 15 s xcrun call on this path ispackages/platform-apple/src/runner/runner-toolchain-probe.ts:160(xcrun --sdk iphonesimulator --show-sdk-version).apple-runner-platform.ts:13already documents the mechanism — on a fresh macOS host Apple's syspolicyd signature verification makes the first xcrun invocation slow, so this is a cold-host cost landing inside a fixed budget.Two things worth deciding here, rather than per-PR:
- Does fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336's re-issue layer absorb this?
COMMAND_FAILEDfrom a preflight step is not obviously a typed-retriable infra miss, so it may red straight through the retry policy that fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336 is building — which would leave every branch red on a cold runner host regardless of the fix. That is a scope question for fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336, not for the PR that happens to trip it. - Is 15 s the right budget for a first-invocation probe on shared CI hosts? I have explicitly instructed the agent on fix(android): keep close --shutdown's IME restore out of the settings flush window #3331 not to widen it or add a retry, because a per-PR widening to paper over a contended host is suppression, not a fix. The budget belongs here.
Not filed as a new issue, and deliberately not folded into #3342: that one is
daemon_startup_failedat the same step, a different failure of the same seam. Keeping the classes separate.Context:
mainis green on iOS across all recent runs, so this is host-lane flake, not a landed regression.- Does fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336's re-issue layer absorb this?
- added a commit that references this issue
on Oct 9, 2026
The fixture-backed iOS simulator E2E smoke job is failing frequently on
main, with different assertions on different runs, which makes it unusable as a merge signal: a red smoke on a PR currently says nothing about that PR.Rate. Of the last 12
ios.ymlruns onmain(2026-09-10 → 2026-09-11): 6 failures, 2 success, 4 cancelled.Signatures observed, all on
mainor reproducing independently of a change:wait for the WebView page to expose its linksmoke:webview-remote-content34592935964id="automation-longpress" did not become visible after scrollingis visible34595714111wait timed out for text: Agent Device Tester,wait_capture_stalled,readableCaptures: 0on the FIRST capture after a coldopen --relaunchsmoke:automation-input34400074702,34226333904— already tracked as #2343xcrun --sdk iphonesimulator --show-sdk-version (spawnSync xcrun ETIMEDOUT)after** TEST BUILD SUCCEEDED **34490275779— matches #2422 (closed, recurring)Test timed out in 5000ms34519750958The clearest evidence that this is the lane and not the change under test: two runs of the same commit (
6ba5b3da81, job103251301238then the re-run103255078739) failed with two different assertions —automation-longpressfirst, then the webview-link one thatmainis also failing. A deterministic regression cannot rotate its symptom.Why it matters now. Reviewers are being asked to judge merge-readiness against a signal that is red for unrelated reasons, and the honest response to each red is a manual bisect against main — which is exactly the cost this lane exists to avoid. #2343 covers one signature; the others are unfiled.
Suggested direction (not prescriptive): step 14 runs
node --testdirectly with no retry layer, while the replay steps in the same job use--retries 2, so a single cold-start stall fails the whole job. Worth deciding per-scenario which failures are genuinely non-retriable, and separating "the lane found a real defect" from "the simulator was cold or slow". The capture-stall class in #2343 already has a measured mechanism to build on.Related: #2343, #2422.