Skip to content

Diagnose recurring multi-minute Settings replay click in iOS Smoke #2948

Description

@thymikee

Purpose

Reduce a recurring multi-minute delay in the iOS Smoke Settings replay and determine whether its successful click General response reflects a completed interaction.

Evidence

Three independent exact-head PR runs on 2026-09-24 show the same pattern in 01-settings.ad: the first attempt reports step 3 click General as successful after 143.840 s, 191.076 s, or 165.128 s; step 4 then times out after 5 s waiting for About, Software Update, or the expected Settings text. The retry succeeds in 30.5 s, 23.6 s, or 14.8 s respectively.

The recorded replay timing establishes the delay and failed wait; it does not establish which internal click phase consumed the time or whether the tap landed. The #2936 failed wait request log includes a simulator-target-discovery-pending bridge fallback, but successful click requests did not retain equivalent phase diagnostics.

Required behavior

  • Attribute the click time to target resolution, snapshot/bridge acquisition, XCTest transport, tap synthesis, and post-action work using existing request/session diagnostics or narrowly scoped new timings.
  • Determine from a live failing attempt whether General was actually activated before reporting click success; repair the owning path if success can be reported without the action, or bound an optional pre-action probe beneath the operation it precedes.
  • Preserve the retry as failure containment while diagnosing; do not treat the retry as proof of click correctness.

Completion

  • A live iOS run captures enough phase evidence to explain the 143–191 s first-attempt clicks and the subsequent failed wait.
  • A focused fix or an evidence-backed infrastructure disposition removes the repeated multi-minute cost without hiding a failed action. Compare exact-head iOS runner time and replay step timing against the linked runs.

Depends on no current PR; this is separate from #2936's depth-frontier diagnostic.

Activity

  1. thymikee commented on Sep 25, 2026

    @thymikee
    MemberAuthor

    The first-click delay is predominantly runner startup, not tap synthesis. In the linked #2936 job, the runner was preflighted (97.0 s), then the workflow ran pnpm clean:daemon, releasing it. The first Settings click took 143.8 s; its runner log shows an interrupted initial launch, a full build-for-testing, another interrupted launch, then a ready runner. Once ready, snapshot and tap completed in under 3 s; the tap was synthesized at (201, 406) and reported success, but the next wait still timed out. That does not prove navigation occurred.

    #2949 keeps the prepared runner alive through the replay. A local dedicated simulator passed the replay in 10.0 s with the same daemon (1.3 s click), versus 12.0 s after cleanup (1.4 s click); the local cache was warm. I am waiting for exact-head CI to measure the cold-run benefit. This issue remains open for the failed post-tap destination check and any remaining startup delay.

  2. thymikee commented on Sep 25, 2026

    @thymikee
    MemberAuthor

    Exact-head #2949 iOS job passed and proves the prepared runner was reused: one preflight xcodebuild launch, runner PID 30721 serving both preflight and Settings commands, and no rebuild before replay. The Settings step fell from 4m50s to 1m28s; first click fell from 143.840 s to 18.189 s.

    The remaining first-attempt failure is real. open reported postOpenObservation: unobservable; first click spent 13.1 s in a runner snapshot after a 4 s system-modal probe timeout and private-AX fallback, then the tap command completed in 1.1 s. The subsequent wait failed after 9.0 s; attempt 2 passed with a 4.7 s click. Runner reuse removes the multi-minute restart/build penalty but does not establish that the first tap navigated. Keep this issue open for the residual snapshot and post-tap behavior.

  3. thymikee commented on Sep 25, 2026

    @thymikee
    MemberAuthor

    Follow-up on the residual after #2949: the failed attempt does not prove a missed tap. The first runner snapshot took 13.09 s (4 s system-modal probe timeout, drain, then private-AX fallback); the tap took 1.07 s. The next wait received no readable capture before its deadline: its bridge acquisition was canceled, and the runner fallback snapshot completed after the wait. The failed attempt tapped (201,319) and the passing retry (201,406), but the older #2936 failure tapped (201,406) too, so coordinate difference is not causal proof.

    A separate local replay passed; its 19.5 s click was cold runner startup and did not reproduce this wait failure. The next diagnostic run should enable global --debug for the Settings test and retain an on-failure screen image or video plus the selected-node rectangle and capture backend. That would tell us whether navigation missed or the wait simply lacked a readable capture. No interaction fix is justified from the current artifacts.

  4. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Phase-attribution report — diagnostic plan executed live, 2026-10-08 (iPhone Duo, iOS 27.1, UDID 4F879835)

    Ran the owning lane's real harness (node src/bin.ts test test/integration/replays/ios/simulator/01-settings.ad --udid … --retries 2, i.e. the ios.yml "Run iOS Settings replay smoke test" command shape, prepare ios-runner first, same shared state dir), 20 replay cycles = 44 attempts plus a cold-daemon no-preflight class and 10+13 focused probe iterations. Evidence: replay-timing.ndjson per attempt, per-request daemon logs (sessions/<s>/requests/<req>.ndjson under --debug + explicit --state-dir), session runner.log, and today's main CI artifact ios-artifacts from run 37815426076 (which reproduced the original signature: attempt 1 failed at step 4, total 57.8s flaky).

    Observed distribution (44 local attempts + 1 live CI failing attempt)

    Phase attribution of the multi-minute click (cold-1 attempt-1, 15.5 s local; same shape as the historic 143–191 s, minus Xcode-cache warmth)

    From the click's own request ndjson: bridge target readiness 0.24 s; runner demand inside the click: build-for-testing 13.1 s (exec_command dur), xctestrun prep 0.4 s, test-without-building launch + port-connect retries 4.9 s (first uptime round-trip dur=4926), readiness preflight, then tap synthesis ios_runner_command_send cmd=tap dur=888 ms. i.e. the multi-minute window is runner build+launch inside the first runner-demanding command — the #2949 preflight-reuse class, already merged; the diagnostic reproduces its shape. Warm click phases: bridge probe/fallback ~0.2 s + readiness snapshot 2.0 s (snapshot_capture dur=2001 backend=xctest) + tap 0.59 s ≈ 2.6 s.

    Did General actually activate before click success?

    • Locally (iOS 27.1): yes, verified. click General succeeded, and the failure-time screen at step 7 is the General detail page (nav title General, cell About); direct probes: 10/10 is exists "label=About" pass immediately after click. The 42 step-7 failures are the back-control defect, not a missed tap. The retry's success here is coincidental floating-toolbar flake, not proof.
    • On the live CI failing attempt: unanswerable from retained evidence, and likely not activated. Attempt 1's tap was synthesized at (201,319) ~0.5 s after a 2.8 s capture of a still-settling root list (that y hits the Apple-Account band); the passing attempt 2 tapped the settled (201,406). The click reported success from the tap ack (tap dur≈0.8 s, ok=1), and the wait then got zero readable captures — its runner snapshot was only COMMAND_ACCEPTED at 17:41:47.2, after the wait had already failed at 17:41:46.4, and completed ok=1 1.6 s later against a dead request. Success was reported without any landing observation. Per the issue's own rule ("file it as a new issue with evidence rather than fixing it here"), that owning-path gap is filed as click and press report dispatch, not landing: state the contract and track the outcomeObservation gap #3335.

    Bridge-circuit correlation test (the #2491/#3328 signature)

    Does not correlate with the multi-minute window. In every local attempt the daemon emitted ios_snapshot_route_fallback (foreground-owner-unverified at open → circuit-disabled for the app generation; 45+96 events), and each fallback cost ≤ ~2.2 s on that capture, never more; clicks stayed at 2.5–3 s throughout. In the CI failing window (17:41–17:42) the replay session's runner log has zero PRIVATE_AX_SNAPSHOT_FAILED / SNAPSHOT_BACKEND_FAILED lines — those signatures in the same job log belong to the runner-XCTest/E2E sessions at 17:24–17:34. The failing wait's 10.4 s against a 5 s budget is a capture that never reached the runner in budget (dispatched after the wait failed), i.e. the observation-stall class of #2491/#3328 — not circuit-consumed time.

    Disposition

    Cleanup: diag2948 session closed, daemon retired via pnpm clean:daemon, iPhone Duo 4F879835 shut down; no other device touched.

  5. removed
    needs-infoWaiting on reporter or external input
    on Oct 8, 2026
  6. thymikee commented on Oct 9, 2026

    @thymikee
    MemberAuthor

    Closing as completed. Each completion clause now has an owner:

    The retry is still in place as containment. No retry success was treated as proof that the click worked.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions