Skip to content

ci(ios): the iOS simulator smoke lane is failing on main across several rotating signatures #2491

Description

@thymikee

The fixture-backed iOS simulator E2E smoke job is failing frequently on main, with different assertions on different runs, which makes it unusable as a merge signal: a red smoke on a PR currently says nothing about that PR.

Rate. Of the last 12 ios.yml runs on main (2026-09-10 → 2026-09-11): 6 failures, 2 success, 4 cancelled.

Signatures observed, all on main or reproducing independently of a change:

assertion scenario / step seen in
wait for the WebView page to expose its link smoke:webview-remote-content main 34592935964
id="automation-longpress" did not become visible after scrolling scroll + is visible PR run 34595714111
wait timed out for text: Agent Device Tester, wait_capture_stalled, readableCaptures: 0 on the FIRST capture after a cold open --relaunch smoke:automation-input main 34400074702, 34226333904 — already tracked as #2343
xcrun --sdk iphonesimulator --show-sdk-version (spawnSync xcrun ETIMEDOUT) after ** TEST BUILD SUCCEEDED ** runner build step PR run 34490275779 — matches #2422 (closed, recurring)
Test timed out in 5000ms "Verify clean-installed Simulator snapshot bridge preparation" PR run 34519750958

The clearest evidence that this is the lane and not the change under test: two runs of the same commit (6ba5b3da81, job 103251301238 then the re-run 103255078739) failed with two different assertions — automation-longpress first, then the webview-link one that main is also failing. A deterministic regression cannot rotate its symptom.

Why it matters now. Reviewers are being asked to judge merge-readiness against a signal that is red for unrelated reasons, and the honest response to each red is a manual bisect against main — which is exactly the cost this lane exists to avoid. #2343 covers one signature; the others are unfiled.

Suggested direction (not prescriptive): step 14 runs node --test directly with no retry layer, while the replay steps in the same job use --retries 2, so a single cold-start stall fails the whole job. Worth deciding per-scenario which failures are genuinely non-retriable, and separating "the lane found a real defect" from "the simulator was cold or slow". The capture-stall class in #2343 already has a measured mechanism to build on.

Related: #2343, #2422.

Activity

  1. thymikee commented on Sep 11, 2026

    @thymikee
    MemberAuthor

    Keep the failures separated by signature. #2493 addresses the WebView landmark probe and the wait that gives up on RUNNER_BUSY; it does not resolve the cold-capture, scrolling or build-timeout cases. Different failures on the same commit justify investigation, but do not establish that every red run is unrelated to a PR. Confirm each current signature against main and keep this issue open for the remaining cases after #2493 lands.

  2. thymikee commented on Sep 11, 2026

    @thymikee
    MemberAuthor

    Root cause for row 1 (wait for the WebView page to expose its link)

    Traced it end to end on the newest occurrence, job 103282380298 (PR #2448 head 28c6e86bee). It is a lane defect, not a change-under-test defect: the failure is the same signature as main's 34592935964, the log has zero system-surface mentions, and the mechanism below is platform-wide and predates the branch.

    The failing step never polled. The error JSON carries none of the wait evidence fields (reason, timeoutMs, polls, captures, waitedMs, readableCaptures). It is the raw capture error:

    command: agent-device wait text Jump to form 20000 ...
    "code": "RUNNER_BUSY",
    "message": "The iOS runner is still finishing a previous command that exceeded its
                execution watchdog (usually an accessibility capture on a heavy or
                animating screen)."
    "hint":    "Wait a few seconds and retry. ..."
    details: { command: "snapshot", lifecycleState: "failed",
               recovery: "runner_reported_failure" }
    

    So: a WebView accessibility capture exceeded the runner's execution watchdog, the next capture got RUNNER_BUSY, and wait gave up on its first poll with 20 s of budget unspent.

    Why it gives up, in three steps.

    1. The runner error is built as retriable. packages/platform-apple/src/runner/runner-session.ts:920 sets retriable: true for RUNNER_BUSY, and the code is in DIAGNOSTIC_ONLY_RUNNER_ERROR_CODES so it stays COMMAND_FAILED with details.runnerErrorCode preserved.

    2. Nothing retries on that. retriable is a wire classification only: request-router.ts and request-finalization.ts copy it outward to the client, and no host-side caller consults it. The only code that reads details.runnerErrorCode for a policy decision is packages/platform-apple/src/alert.ts:149, and it reads ALERT_NOT_FOUND. RUNNER_BUSY appears in that file and in alert-contract.ts only as prose precedent for the pattern.

    3. The wait poll loop has exactly one error-tolerance channel, and it is Android-only. wait text builds its polling at src/commands/interaction/runtime/wait-text.ts:19 with no classification argument, so createWaitPolling defaults to isUnreadableCaptureContentError, which inspects details.androidSnapshotHelperFailureReason and nothing else. Any other throw propagates out of unreadable.attempt and out of the whole wait.

    The wait family therefore has a budget it cannot spend on the one error whose own hint says to wait and retry. The lane is a heavy animating WebView, which is precisely the screen that trips the watchdog, so this row will keep rotating back until the tolerance channel admits a retriable runner stall.

    Filed as #2496 for the fix. Recording it here so row 1 has a cause rather than a count.

  3. thymikee commented on Sep 11, 2026

    @thymikee
    MemberAuthor

    Correction to my comment above, and the fix is already in flight: #2493.

    Two things in my diagnosis were wrong or incomplete.

    The trigger is not "a heavy animating WebView trips the watchdog." It is a landmark mismatch. acceptDeepLinkConfirmationIfPresent had a hard-coded Automation-lab readiness landmark, so on any other route it could never match and always fell through to its alert get probe. An XCTest alert query against a live WKWebView exceeds the runner's 30 s main-thread execution watchdog, and that is what leaves the runner refusing later commands as RUNNER_BUSY. smoke:regular-visible-depth-frontier carried the same mismatch and the same doomed probe. So row 1 is a scenario-authoring defect first, and the wait's surrender is what turned it into an opaque failure.

    There is a second classification path I did not find. A runner error recovered from the lifecycle journal after a lost transport response was built with a bare toAppErrorCode, so on that path RUNNER_BUSY reached callers as a wire code with no retriable flag at all.

    #2493 fixes all three: the route landmark, the wait ride-out, and one classifyRunnerReportedError behind both paths. It carries live before/after validation on a booted simulator. My wait-poll trace above is accurate as a mechanism but is superseded as a root cause. I also filed #2496 before finding #2493 and have closed it as a duplicate.

    Row 1 should be considered owned by #2493 rather than open here.

  4. thymikee commented on Sep 11, 2026

    @thymikee
    MemberAuthor

    Keep row 1 assigned to #2493, but not resolved yet: its latest exact-head iOS smoke still fails at the WebView page wait. The error now has the corrected COMMAND_FAILED / retriable classification, which confirms that part of the fix reached the lane; successful observation of the page is still missing. The other signatures remain separate work.

  5. okwasniewski commented on Sep 23, 2026

    @okwasniewski
    Contributor

    Today's signature on the iOS smoke lane, hitting every PR I have open plus two unrelated ones:

    AssertionError: regular depth-1 snapshot must disclose the Simulator AX bridge evidence gap
      at assertSimulatorBridgeSnapshot (test/integration/ios-simulator-e2e/live-snapshot-depth-frontier.ts:126)
    

    Runs: 35862236341 (#2804), 35862934402 (#2811), 35862504224 (#2814), 35862792165 (fix/scroll-movement-disclosure), 35845257256 (fix/native-stack-bridge-hittability). The three PRs of mine touch the daemon resend policy, the daemon-client takeover decision, and two runner log lines respectively; none change the snapshot backend plan.

    In the failing runs the regular depth-1 snapshot came back from the XCTest tree (snapshotDiagnostics.backends: {xctest: 54}) with the slow-snapshot warning (p95 3426ms, max 5626ms), so the bridge tier is not the one answering under that load, and the assertion that expects bridge evidence fails. Looks like host load on the runners rather than a code regression; the same job passes on rerun (#2804 had one pass and one fail on the same head).

  6. thymikee commented on Sep 24, 2026

    @thymikee
    MemberAuthor

    Status after a census of 123 failed iOS Smoke jobs (09-14 to 09-23; main red on 17 of 59 runs). Each rotating signature now has a named mechanism and a merged fix:

    Signature Failures Fix
    must disclose the Simulator AX bridge evidence gap 22 (all on 09-23) #2832: #2775 had merged with a stale E2E assertion
    first wait for Agent Device Tester after open --relaunch (wait_capture_stalled) 15 #2838: open joins the in-flight discovery
    wait for Automation lab after the deep-link "Open in…" prompt 12 #2852: runner reads never launch a stopped session app
    automation-longpress did not become visible 20 #2839: reads after an unsettled gesture carry unsettledGesture and the helper re-reads
    alert-observation XCTests 13 #2843 (closes #2546)
    testAbandonedTreeCapture… 8 #2847

    Still open or unaddressed: the webview-link wait (1 sighting in the window), cold app.launch() timeouts in the XCTest lane (4), #2343 (wait startup accounting), and the follow-ups #2853, #2856 and #2862. I'm keeping this open until main shows a run of consecutive green iOS Smoke results. That is also when iOS Smoke could become a required check, since #2775 merged red.

  7. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Fresh attribution pass (2026-10-08, commits a513106f7…24b2ce638)

    I pulled the failing-step payloads for every recent ios.yml main failure (gh run view <id> --repo callstack/agent-device --json jobs, then gh run download <id> --repo callstack/agent-device -n "ios-e2e-simulator-summary-*" / -n "ios-e2e-simulator-replay-*" and read work/ios-steps/<scenario>/failed-step.txt). All 10 recent main failures now resolve to known signatures plus three new ones:

    Run Signature Attribution
    37459193672, 37145696803 snapshotQuality warning on snapshot --base64 #3328 — misasserted test (product always stamps a strategy on the circuit-disabled XCTest fallback; the AX bridge publishes no verdict). Fix: #3336
    37419258095, 37187537985 Settings replay: wait text Automation lab deadline exceeded, @preset-detail not found, attempt 3/3, step 4/11 The #2948-adjacent Settings replay wait — recorded here so it is no longer unattributed (owned under the #2948 follow-up work)
    37356199982 (wait_runner_restart_exhausted), 36454269522 (wait_capture_stalled), 36397945147 + 36247999237 (wait_readiness_exhausted), 36298519121 (wait_readiness_exhausted, runner restart mid-run) hard failure on wait text "Automation lab" in smoke:automation-input E2E step Infra-flaked transport/observation waits — the class the retry policy in #3336 now re-issues once
    36115336398 gesture-pan replay TIMEOUT + timeout_cleanup_pending (runner 310s, wall 452s) New, untracked — replay-timeout class, cleanup pending after command timeout
    36127405340 xcrun --sdk iphoneos --show-sdk-version ETIMEDOUT during preflight Tracked: #2422
    37660079219 alert dismiss → MAIN_THREAD_TIMEOUT (phase after_command_dispatch) on smoke:interaction-registry step 10/17 New, untracked — #2782 (finished-at-boundary misreport) is closed and does not cover this; looks like a genuine main-thread stall during the interaction-registry scenario; adjacent to #2956/#2546 alert flakiness
    36916418541, 36472317113 run cancelled mid-E2E with wait_runner_restart_exhausted (UI-interaction scenario) at retriable: true while the job was observation-prevented Not a scenario verdict — cancelled runs whose last step carried a retriable wait reason; same transport/observation class as the row above

    Two earlier entries in the sample (37043315673, 37033991194) turned out to be pull_request-event runs, not push-event main runs — excluded.

    Measured failure rate

    Sample: last 200 ios.yml runs on main, window 2026-09-25 → 2026-10-08.

    gh run list --repo callstack/agent-device --workflow ios.yml --branch main --limit 200 \
      --json databaseId,conclusion,event,createdAt
    gh run list --repo callstack/agent-device --workflow "iOS Backward Compatibility" --branch main --limit 100 \
      --json databaseId,conclusion,event,createdAt
    • 191 push-event ios.yml main runs: 16 failures / 64 successes / 111 cancelled → failure rate among completed push runs ≈ 20% (16/80).
    • Monthly split: September 9 fails / 32 completed ≈ 28%; October 7 / 48 ≈ 15%.
    • Backward-compat lane: all non-failing (32 successes, 66 cancelled).

    The cancellations are dominated by concurrency-group supersedes (new pushes), so the meaningful denominator is completed runs only.

    Retry policy decision (#3336)

    The lane's two replay steps already run under --retries 2 (node test runner); the E2E steps had no layer, so every transport/observation miss ruled red immediately. #3336 adds a shared, opt-in harness seam (reattemptInfrastructureMiss in live-device-e2e/runtime.ts) that re-issues one failed step, enabled by the iOS lane only when the miss is typed as infrastructure: error.retriable === true AND wait reason ∈ {wait_capture_stalled, wait_runner_restart_exhausted, wait_readiness_exhausted}. No error-text matching.

    Non-retriable (stays red on first failure): wait_target_absent, wait_deadline_exceeded, wrong asserted values, runner crashes, and any non-wait failure; allowFailure/expectFailure steps are exempt. Retried misses still print their first failure before the re-issue; the lane summary reports the count. Bounded at one re-issue per step, INFRASTRUCTURE_MISS_ISSUES = 2 in the runtime.

    Known limitation: the re-issue shares the same runner/session, so a runner-dead miss can re-fail unchanged; per-scenario fresh-runner teardown is the lane-rewrite follow-up this feeds into.

  8. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Correction to the table above (caught re-checking the downloaded payloads before filing #3337):

    • 37660079219: the MAIN_THREAD_TIMEOUT on alert dismiss occurred in smoke:automation-input (step 55, "dismiss native alert"), not smoke:interaction-registry. Typed details: reason: runner_main_thread_timeout, dispatched: "unknown", runnerFailureReason: runner_main_thread_execution_timeout. The failing run's step history contains only bootstrap, smoke:inventory-install, and smoke:automation-input.
    • 36115336398: the failing step was "Run gesture pan-duration smoke replay" — examples/test-app/replays/gesture-pan-duration.ad, attempt 1/3, TIMEOUT after 180000ms on a wait command with timeoutCleanupPending: true, junit wall time 182s, and a subsequent ** BUILD INTERRUPTED ** in that attempt's runner log. The settings replay in the same run passed (7 steps replayed). The "runner 310s / wall 452s" and "12 pass / 6 fail" figures quoted in the original row do not exist in this run's artifacts — disregard them.

    Everything else in the comment stands; both corrections are reflected in #3337.

  9. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Ledger update: a fourth untracked signature appeared on a PR lane (not main) — smoke:webview-remote-content cold-start overrun, wait_deadline_exceeded on wait text "Jump to form" 20000 with the page content present in the post-failure snapshot (run 37843633972, PR #3336 lane). Filed as #3343. Also filed: #3342 for the preflight daemon_startup_failed PR-lane class (two branches, identical signature; did not reproduce on re-run at 98321a2). The #3336 lane-policy intentionally leaves both red until their owners land.

  10. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Ledger correction to my previous comment: the webview cold-start signature (wait_deadline_exceeded with high readableCaptures, content present in the post-failure snapshot) is now recorded in #3337 as its third scenario-level signature (per the #3336 coordinator pass); my standalone #3343 is closed as a duplicate with a forwarding pointer. The preflight daemon_startup_failed class stays in #3342, now corroborated by a third cross-branch instance (#3331 run 37833483850 attempt 1, 19:44Z) and non-reproduction on head 98321a2.

  11. thymikee commented on Oct 9, 2026

    @thymikee
    MemberAuthor

    Attribution update from the PR-supervision coordinator: this signature now also reaches the Preflight iOS runner through public CLI step, where it fails the job before any test runs.

    Instance (2026-10-09T00:40:20Z, attempt 1): run 37865442839, job 113611128846, head ce40232e3 on fix/android-shutdown-ime-flush-window (PR #3331, an Android-only diff — it cannot touch an Apple path). Failed step: Preflight iOS runner through public CLI. Payload:

    "code": "COMMAND_FAILED"
    "message": "xcrun timed out after 15000ms"
    "hint": "Retry with --debug and inspect diagnostics log for details."
    "diagnosticId": "mv08mt9n-400b51b8"
    

    Same signature, new location. This table already records xcrun --sdk iphonesimulator --show-sdk-version (spawnSync xcrun ETIMEDOUT) after ** TEST BUILD SUCCEEDED ** in the runner-build step, matching #2422. This time the identical probe call is what fails, but it runs earlier: the preflight step is the first prepare ios-runner on the lane, so the failure aborts the job with zero smoke steps executed and no step-history.json, which is why it reads as a different class than it is.

    Call site, for whoever owns the budget: exec.ts:368 builds the timed out after <n>ms message, and the only 15 s xcrun call on this path is packages/platform-apple/src/runner/runner-toolchain-probe.ts:160 (xcrun --sdk iphonesimulator --show-sdk-version). apple-runner-platform.ts:13 already documents the mechanism — on a fresh macOS host Apple's syspolicyd signature verification makes the first xcrun invocation slow, so this is a cold-host cost landing inside a fixed budget.

    Two things worth deciding here, rather than per-PR:

    1. Does fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336's re-issue layer absorb this? COMMAND_FAILED from a preflight step is not obviously a typed-retriable infra miss, so it may red straight through the retry policy that fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336 is building — which would leave every branch red on a cold runner host regardless of the fix. That is a scope question for fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) #3336, not for the PR that happens to trip it.
    2. Is 15 s the right budget for a first-invocation probe on shared CI hosts? I have explicitly instructed the agent on fix(android): keep close --shutdown's IME restore out of the settings flush window #3331 not to widen it or add a retry, because a per-PR widening to paper over a contended host is suppression, not a fix. The budget belongs here.

    Not filed as a new issue, and deliberately not folded into #3342: that one is daemon_startup_failed at the same step, a different failure of the same seam. Keeping the classes separate.

    Context: main is green on iOS across all recent runs, so this is host-lane flake, not a landed regression.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions