Problem
Android MDM commands (Lock, Wipe, Clear passcode) issued by Fleet via AMAPI EnterprisesDevicesService.IssueCommand are tracked in mdm_android_commands and start with status='pending'. The lifecycle relies on the Pub/Sub COMMAND notification from AMAPI to advance the row to acknowledged (success) or error (failure), and to clear the matching host_mdm_actions.lock_ref / wipe_ref. If that Pub/Sub notification never arrives (for example because Fleet's Pub/Sub push endpoint is unavailable for longer than GCP's 7-day retention, or because of a subscription misconfiguration or Google Cloud incident), the row stays pending forever. host_mdm_actions keeps the ref set, lockWipe.IsPendingLock() / IsPendingWipe() return true forever, and the host shows a perpetually-pending state in any UI driven by those checks. The admin cannot re-issue the command because Fleet's pending-state guard blocks it.
Apple and Windows MDM have the same failure shape today: a dropped CheckIn response (Apple) or sync session result (Windows) leaves nano_command_results / windows_mdm_command_results empty and Fleet's state stuck "pending." Fleet does not currently reconcile either platform. Android is different in one important way that makes this fix tractable: AMAPI exposes a direct operations.get(operation_name) REST endpoint that returns the authoritative command state, so Fleet can recover without depending on the device ever coming back online or AMAPI re-sending a notification.
Impact
Severity: low. GCP Pub/Sub provides at-least-once delivery with retries up to 7 days, so drops in normal operation are uncommon. The most realistic failure mode is a Fleet endpoint outage longer than 7 days that overlaps an in-flight Android command. Frequency: rare in practice, but unbounded once it happens — affected rows never self-recover. Blast radius: per-host, per-command. The admin loses the ability to re-issue Lock/Wipe/Clear-passcode on an affected host until someone manually clears the stale row in MySQL. Same conceptual failure exists on Apple and Windows already, so this proposal does not regress platform parity; it brings Android forward.
Proposed fix
Add a Fleet cron (suggested name mdm_android_command_reconcile) that runs daily. The cron scans mdm_android_commands WHERE status='pending' AND created_at < NOW(6) - INTERVAL 1 DAY in batches. For each row, call AMAPI operations.get(operation_name) to fetch the authoritative state. If Operation.Done == true, update the row's status (and error_code / error_message from Operation.Error if set) and clear the matching host_mdm_actions.lock_ref or wipe_ref so the host returns to a re-issuable state. If Operation.Done == false, leave the row alone — the command is still queued at AMAPI. Rate-limit AMAPI calls (e.g., 50 ops/minute, with backoff on quota errors) to stay well under AMAPI's per-project request budget. Daily cadence is sufficient because the failure mode is rare and a 24-hour reconciliation lag is invisible to admins who would otherwise wait indefinitely.
Extending the same mechanism to Apple and Windows is a separate question. Apple has APNS-push-on-demand and Windows has sync session triggers; both are more involved than calling a single REST endpoint, so they belong in their own follow-up issues.
Evidence
PR #46031 review by ksykulev on 2026-05-23: #46031 (review) — explicitly raised the dropped-notification concern and asked whether there are plans to handle it.
The Operation-resource state returned by operations.get is the authoritative reconciliation source.
Problem
Android MDM commands (Lock, Wipe, Clear passcode) issued by Fleet via AMAPI
EnterprisesDevicesService.IssueCommandare tracked inmdm_android_commandsand start withstatus='pending'. The lifecycle relies on the Pub/SubCOMMANDnotification from AMAPI to advance the row toacknowledged(success) orerror(failure), and to clear the matchinghost_mdm_actions.lock_ref/wipe_ref. If that Pub/Sub notification never arrives (for example because Fleet's Pub/Sub push endpoint is unavailable for longer than GCP's 7-day retention, or because of a subscription misconfiguration or Google Cloud incident), the row stayspendingforever.host_mdm_actionskeeps the ref set,lockWipe.IsPendingLock()/IsPendingWipe()return true forever, and the host shows a perpetually-pending state in any UI driven by those checks. The admin cannot re-issue the command because Fleet's pending-state guard blocks it.Apple and Windows MDM have the same failure shape today: a dropped CheckIn response (Apple) or sync session result (Windows) leaves
nano_command_results/windows_mdm_command_resultsempty and Fleet's state stuck "pending." Fleet does not currently reconcile either platform. Android is different in one important way that makes this fix tractable: AMAPI exposes a directoperations.get(operation_name)REST endpoint that returns the authoritative command state, so Fleet can recover without depending on the device ever coming back online or AMAPI re-sending a notification.Impact
Severity: low. GCP Pub/Sub provides at-least-once delivery with retries up to 7 days, so drops in normal operation are uncommon. The most realistic failure mode is a Fleet endpoint outage longer than 7 days that overlaps an in-flight Android command. Frequency: rare in practice, but unbounded once it happens — affected rows never self-recover. Blast radius: per-host, per-command. The admin loses the ability to re-issue Lock/Wipe/Clear-passcode on an affected host until someone manually clears the stale row in MySQL. Same conceptual failure exists on Apple and Windows already, so this proposal does not regress platform parity; it brings Android forward.
Proposed fix
Add a Fleet cron (suggested name
mdm_android_command_reconcile) that runs daily. The cron scansmdm_android_commands WHERE status='pending' AND created_at < NOW(6) - INTERVAL 1 DAYin batches. For each row, call AMAPIoperations.get(operation_name)to fetch the authoritative state. IfOperation.Done == true, update the row'sstatus(anderror_code/error_messagefromOperation.Errorif set) and clear the matchinghost_mdm_actions.lock_reforwipe_refso the host returns to a re-issuable state. IfOperation.Done == false, leave the row alone — the command is still queued at AMAPI. Rate-limit AMAPI calls (e.g., 50 ops/minute, with backoff on quota errors) to stay well under AMAPI's per-project request budget. Daily cadence is sufficient because the failure mode is rare and a 24-hour reconciliation lag is invisible to admins who would otherwise wait indefinitely.Extending the same mechanism to Apple and Windows is a separate question. Apple has APNS-push-on-demand and Windows has sync session triggers; both are more involved than calling a single REST endpoint, so they belong in their own follow-up issues.
Evidence
PR #46031 review by ksykulev on 2026-05-23: #46031 (review) — explicitly raised the dropped-notification concern and asked whether there are plans to handle it.
The Operation-resource state returned by
operations.getis the authoritative reconciliation source.