Skip to content

[Phase 2] Add scheduled-run reliability analysis to the fleet health daily check #725

Description

@github-actions

Story

As a fleet reliability owner,
I want extend the daily Actions Fleet Monitor to measure and report per-workflow scheduled-run reliability -- expected scheduled ticks vs. actual schedule-triggered runs in the window,
so that a cron that silently never fires becomes visible and distinguishable from a healthy one, so 'successful but never ran' stops hiding stalled backlogs.

Acceptance Criteria

  1. For each workflow that has a schedule trigger, fleet_monitor.sh/fleet_report.sh compute the number of expected scheduled ticks within the lookback window and the number of actual runs whose triggering event is schedule, plus a reliability/hit-rate metric (and/or a missed-tick count).
  2. The per-workflow runs query captures the run's event field so schedule-triggered runs are isolated from workflow_dispatch / push / repository_dispatch runs.
  3. The workflow's cron expression(s) are obtained per workflow (parsed from the workflow file YAML, since the actions/workflows API does not expose cron) and expanded to the count of expected ticks over the window; multi-cron workflows and the default-branch-only scheduling caveat are handled or explicitly documented.
  4. The daily fleet report (both the GITHUB_STEP_SUMMARY and the issue-body report) gains a schedule-reliability section, and workflows whose reliability falls below a threshold are surfaced in a style consistent with the existing high-failure tracking.
  5. Best-effort delay is not falsely reported as a miss: a tolerance window is applied (a run that fired late but within tolerance counts as present), and the existing caveat that GitHub caps run queries at 1000 results per window is respected/noted.
  6. bats coverage is added (in tests/fleet_report.bats or a new bats file appended to the lint.yml bats list) for the cron-to-expected-ticks expansion and the reliability classification, driven by fixture run data.

Tasks / Subtasks

Dev Notes

  • Entry points: fleet_monitor.sh fetches per-workflow runs (~lines 100-120) selecting run_number/conclusion/created_at/html_url/duration_s -- it does NOT currently request .event; add it there. Metrics are emitted as a 12-field TSV (~lines 154-160) consumed by fleet_report.sh generate_report; extend the row or emit a parallel schedule-metrics file rather than breaking the existing 12-field contract relied on by the high-failure JSON export (~lines 227-240).
  • The actions/workflows API (used at ~line 85) returns id/path/state but NOT the cron schedule -- cron must be parsed from the workflow YAML. For org-wide scanning that means fetching file content per workflow; keep it bounded and tolerate parse failures (record as unknown rather than failing the monitor, consistent with the existing ERROR-sentinel pattern).
  • Scheduled events are best-effort by design (the whole point of Improving reliability of scheduled GitHub Actions runs (cron is best-effort) #656): a late-but-present run must NOT be scored as a miss. Apply a tolerance window and only flag workflows where actual scheduled runs fall materially below expected. Respect the documented 1000-results-per-window cap (comment near line 98).
  • 'Fleet health daily check' = the Actions Fleet Monitor (actions-fleet-monitor.yml, daily) driven by scripts/fleet_monitor.sh + scripts/fleet_report.sh -- that is the surface the owner asked to incorporate this into, not daily-pr-review-health.yml.
  • Testing standard: this repo uses bats with an explicit allow-list in lint.yml; tests/fleet_report.bats already exists and is the natural home for report-rendering and math assertions. python3 is available in CI (lint.yml installs pyyaml/jsonschema) if a small parser helper is cleaner than pure bash/awk.

Project Structure Notes

Changes are confined to scripts/fleet_monitor.sh, scripts/fleet_report.sh, and tests (tests/fleet_report.bats or a new bats file added to lint.yml's bats list). actions-fleet-monitor.yml itself likely needs no change since it just runs fleet_monitor.sh; confirm no new step is required to surface the section.

References

  • scripts/fleet_monitor.sh#per-workflow-runs-query
  • scripts/fleet_monitor.sh#metrics-tsv
  • scripts/fleet_monitor.sh#high-failure-json-export
  • scripts/fleet_report.sh#generate_report
  • .github/workflows/actions-fleet-monitor.yml
  • tests/fleet_report.bats
  • .github/workflows/lint.yml#bats

Likely target surface

  • scripts/fleet_monitor.sh
  • scripts/fleet_report.sh
  • tests/fleet_report.bats
  • .github/workflows/lint.yml

Story prepared by the BMAD Scrum Master (Bob) for epic #722. Status: ready-for-dev.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dev-leadFor dev-lead agent pickupinitiativeEpic / initiative tracking issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions