Skip to content

Android MDM command reconciliation via AMAPI operations.get polling cron #46145

Description

@getvictor

Problem

Android MDM commands (Lock, Wipe, Clear passcode) issued by Fleet via AMAPI EnterprisesDevicesService.IssueCommand are tracked in mdm_android_commands and start with status='pending'. The lifecycle relies on the Pub/Sub COMMAND notification from AMAPI to advance the row to acknowledged (success) or error (failure), and to clear the matching host_mdm_actions.lock_ref / wipe_ref. If that Pub/Sub notification never arrives (for example because Fleet's Pub/Sub push endpoint is unavailable for longer than GCP's 7-day retention, or because of a subscription misconfiguration or Google Cloud incident), the row stays pending forever. host_mdm_actions keeps the ref set, lockWipe.IsPendingLock() / IsPendingWipe() return true forever, and the host shows a perpetually-pending state in any UI driven by those checks. The admin cannot re-issue the command because Fleet's pending-state guard blocks it.

Apple and Windows MDM have the same failure shape today: a dropped CheckIn response (Apple) or sync session result (Windows) leaves nano_command_results / windows_mdm_command_results empty and Fleet's state stuck "pending." Fleet does not currently reconcile either platform. Android is different in one important way that makes this fix tractable: AMAPI exposes a direct operations.get(operation_name) REST endpoint that returns the authoritative command state, so Fleet can recover without depending on the device ever coming back online or AMAPI re-sending a notification.

Impact

Severity: low. GCP Pub/Sub provides at-least-once delivery with retries up to 7 days, so drops in normal operation are uncommon. The most realistic failure mode is a Fleet endpoint outage longer than 7 days that overlaps an in-flight Android command. Frequency: rare in practice, but unbounded once it happens — affected rows never self-recover. Blast radius: per-host, per-command. The admin loses the ability to re-issue Lock/Wipe/Clear-passcode on an affected host until someone manually clears the stale row in MySQL. Same conceptual failure exists on Apple and Windows already, so this proposal does not regress platform parity; it brings Android forward.

Proposed fix

Add a Fleet cron (suggested name mdm_android_command_reconcile) that runs daily. The cron scans mdm_android_commands WHERE status='pending' AND created_at < NOW(6) - INTERVAL 1 DAY in batches. For each row, call AMAPI operations.get(operation_name) to fetch the authoritative state. If Operation.Done == true, update the row's status (and error_code / error_message from Operation.Error if set) and clear the matching host_mdm_actions.lock_ref or wipe_ref so the host returns to a re-issuable state. If Operation.Done == false, leave the row alone — the command is still queued at AMAPI. Rate-limit AMAPI calls (e.g., 50 ops/minute, with backoff on quota errors) to stay well under AMAPI's per-project request budget. Daily cadence is sufficient because the failure mode is rare and a 24-hour reconciliation lag is invisible to admins who would otherwise wait indefinitely.

Extending the same mechanism to Apple and Windows is a separate question. Apple has APNS-push-on-demand and Windows has sync session triggers; both are more involved than calling a single REST endpoint, so they belong in their own follow-up issues.

Evidence

PR #46031 review by ksykulev on 2026-05-23: #46031 (review) — explicitly raised the dropped-notification concern and asked whether there are plans to handle it.

The Operation-resource state returned by operations.get is the authoritative reconciliation source.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

#g-byodProduct group focused on Android BYODreliability

Type

No type

Projects

Milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions