Skip to content

Android: durable flush evidence so a daemon restart inside the test-IME flush window cannot release adb emu kill #3346

Description

@thymikee

Purpose

A confirmed Android test-IME restore currently owes a ~2.5 s SettingsProvider flush hold (MAX_WRITE_SETTINGS_DELAY_MILLIS + margin) before any adb emu kill, but the deadline lives only in the daemon's in-process map (testImeLastRestoreAtPerfMs, packages/platform-android/src/ime-state.ts). If the daemon restarts inside the window, the new process has no record of the recent write while the durable recovery marker has already been cleared (the write is confirmed, so the marker has done its job). A kill-bound close --shutdown or standalone shutdown in the new process then waits for nothing and can kill mid-flush: the provider's delayed XML write is lost and the emulator reboots on the test IME — the user-visible outcome of #3318, arriving by the restart route.

PR #3331 landed the entire in-process half (registration beside the confirming write, kill-site choke-point hold, monotonic-clock timing, mid-wait re-read) and explicitly documented this restart boundary in the ime-state.ts map comment. This issue is the cross-restart half.

Required behavior

  1. After a confirmed ime set restore on an emulator, a daemon restart before the flush window elapses must not allow adb emu kill to run before the newest write's window has passed — evaluated from evidence that survives process death.
  2. Clock constraint. The deadline cannot be persisted in the current shape: the in-process window is keyed off performance.now() deliberately (a forward wall-clock step must not shorten it; fix(android): keep close --shutdown's IME restore out of the settings flush window #3331 has a regression test proving that), and performance.now() is relative to process start, so the number is meaningless to the next process. Any persisted form must be one of (non-exhaustive):
    • an epoch-ms deadline plus a documented rule for what host-clock skew can and cannot do to it (bounded, fail-safe toward waiting); or
    • a monotonic-safe derivation (e.g. boot-identifying token plus elapsed rule); or
    • a durable "flush unconfirmed" marker that a kill-bound path must re-verify against the device rather than a host clock.
      State the chosen form and its skew behavior in the implementation.
  3. Expiry rule is mandatory. Durable flush evidence without an expiry is worse than today's hole: a stale marker could block every future shutdown of a dead/removed AVD forever. Evidence must become inert once the window has provably passed (wall time past deadline ⇒ the provider flush completed long ago ⇒ delete-on-read or equivalent), and a vanished/unkillable device must not strand the shutdown path — decide explicitly whether expiry skips the wait or the kill proceeds with a warning.
  4. Startup must keep paying zero sleep for the window: restore-time registration stays synchronous; only kill-bound paths consume it (the invariant fix(android): keep close --shutdown's IME restore out of the settings flush window #3331 encoded and tested).

Observable completion conditions

  • Unit/kill-site tests at the level of packages/platform-android/src/shutdown/runtime.test.ts: a kill-bound close in a "fresh process" (test resets the in-process map but not the durable evidence) waits out the persisted window before adb emu kill; the deferred-sleep interleaving discipline already established there applies.
  • A test showing expired durable evidence does not block a kill and self-removes.
  • A test pinning the clock choice against the skew direction it tolerates, mirroring the monotonic-jump test fix(android): keep close --shutdown's IME restore out of the settings flush window #3331 added.
  • No new sleep on daemon startup or ordinary close paths.

Dependencies

  • Lands on the mechanism from PR fix(android): keep close --shutdown's IME restore out of the settings flush window #3331 (registration inside restoreAndroidTestImeFor, kill-site hold in shutdown/runtime.ts, single marker-clear rule via isRecoveryMarkerClearEarned). Do not re-implement the in-process window; extend its evidence.
  • Likely touches the marker format (ime-recovery-marker.ts, currently {version: 1, serial}) and possibly threads a state-dir file host through the shutdown-runtime contract (today Pick<DeviceShutdownRuntimeDependencies, 'commands'>), which is a contracts change to weigh during design review.
  • The cheapest acceptable variant — re-register a full window at startup whenever startup recovery clears a marker after a confirmed restore (accepted over-wait, monotonic-safe, no persistence) — closes only the startup-recovery route, not the close-time route; if chosen, say which routes remain open.

Origin: cubic P1 on PR #3331 (thread discussion_r4225418232), confirmed and analyzed there with the monotonic-vs-persist tension spelled out.

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions