Skip to content

[Phase 4] Wire opt-in remediation into the fleet-monitor workflow behind a human go/no-go #1152

Description

@github-actions

Story

As a fleet platform maintainer,
I want add an opt-in, DRY_RUN-default remediation step to the fleet monitor (after drift detection) with a workflow_dispatch input to go live for the pilot repo, and record the go/no-go decision,
so that remediation runs on the same cadence as detection without ever going live by accident.

Acceptance Criteria

  1. actions-fleet-monitor.yml gains an opt-in remediation step (or a sibling workflow) that runs scripts/fleet_stub_remediate.sh AFTER fleet_stub_drift.json is produced, consuming that artifact.
  2. The step defaults to DRY_RUN=true; going live requires an explicit workflow_dispatch input (e.g. remediate=live) AND the Phase-3 pilot scope — scheduled monitor runs never remediate live.
  3. The step runs under the credential confirmed in the epic's untracked prerequisite to have cross-repo contents:write + pull_requests:write; if that credential is absent, the step degrades to dry-run with a ::warning:: and never fails the monitor.
  4. The human go/no-go for enabling live pilot remediation (and, later, fleet-wide) is recorded (issue/PR/discussion comment), consistent with the repo's go/no-go pattern.
  5. AGENTS.md documents the remediation loop: the allowlist, the pilot gate, the per-run cap, and the DRY_RUN default.
  6. Existing workflow lint passes; the change edits the org-private monitor workflow (not a thin-caller stub), so the thin-caller no-edit constraint is not violated.

Tasks / Subtasks

Dev Notes

  • actions-fleet-monitor.yml already runs fleet_monitor.sh, emits fleet_stub_drift.json, sets HAS_STUB_DRIFT, and dispatches dev-lead on drift — add the remediation step in that same job after the existing report/dispatch steps so it consumes the already-written artifact. [Source: .github/workflows/actions-fleet-monitor.yml]
  • This workflow is org-private automation, NOT a thin-caller stub, so it may be extended (unlike claude.yml / agent-shield.yml / auto-rebase.yml). [Source: CLAUDE.md, AGENTS.md]
  • DRY_RUN-default + explicit-dispatch-to-go-live matches the repo's dry-run-first canary pattern; the credential fallback-to-warning matches fleet_monitor.sh's token-degradation posture (::warning:: instead of failing). [Source: scripts/fleet_monitor.sh, AGENTS.md]
  • The go/no-go record is the human gate before widening from pilot to fleet-wide; keep the epic inert until that decision is made (consistent with the initiative-planner inert-epic model).

Project Structure Notes

Edits the org-private .github/workflows/actions-fleet-monitor.yml plus an AGENTS.md docs section; no new scripts.

References

  • .github/workflows/actions-fleet-monitor.yml
  • scripts/fleet_monitor.sh
  • AGENTS.md
  • CLAUDE.md

Likely target surface

  • .github/workflows/actions-fleet-monitor.yml
  • AGENTS.md

Story prepared by the BMAD Scrum Master (Bob) for epic #1148. Status: ready-for-dev.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dev-leadFor dev-lead agent pickupinitiativeEpic / initiative tracking issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions