Problem: Two disconnected observability pipelines exist for this self-host instance (edge-nl-01). Prometheus/Alertmanager has 34 alert rules and genuinely routes to a live, watched Discord channel. Sentry independently captures rich, well-contextualized error data — but its only notification path is the shipped default alert rule (email to issue owners, falling through to active members), and every one of its 28 currently-open issues is unassigned. Neither pipeline reaches PagerDuty. The result: real, ongoing, high-volume production bugs (one at 4,030 events over 12 days) went completely unnoticed until a manual audit found them.
Area: Self-host observability / alerting
Proposal: Consolidate both pipelines into one real notification path an operator actually sees in real time, tracked as three independently-shippable sub-issues.
Deliverables:
- Sub-issues filed and linked below, each independently mergeable:
- After all three: a real problem in either Sentry or Alertmanager reaches a human within minutes, not via an unread inbox.
Acceptance criteria:
Boundaries:
Cross-reference (added 2026-07-12)
See also the new AMS observability batch in milestone #25 (Miner Wave 4 — AMS Hardening & Packaging): #5184 (Grafana datasource for AMS's local ledgers), #5185 (coding-agent usage/cost dashboard), and #5186-#5188 (Prometheus alert rules for AMS). A self-hoster running both ORB and AMS on one box should get ONE consolidated alerting pipeline, not two independently-built ones — keep this epic's design in mind when AMS's alert rules land.
Problem: Two disconnected observability pipelines exist for this self-host instance (edge-nl-01). Prometheus/Alertmanager has 34 alert rules and genuinely routes to a live, watched Discord channel. Sentry independently captures rich, well-contextualized error data — but its only notification path is the shipped default alert rule (email to issue owners, falling through to active members), and every one of its 28 currently-open issues is unassigned. Neither pipeline reaches PagerDuty. The result: real, ongoing, high-volume production bugs (one at 4,030 events over 12 days) went completely unnoticed until a manual audit found them.
Area: Self-host observability / alerting
Proposal: Consolidate both pipelines into one real notification path an operator actually sees in real time, tracked as three independently-shippable sub-issues.
Deliverables:
severity: criticalAcceptance criteria:
severity: criticalPrometheus alert reaches PagerDuty, not just Discord. (Add a native PagerDuty receiver to Alertmanager for severity:critical alerts #5009)Boundaries:
Cross-reference (added 2026-07-12)
See also the new AMS observability batch in milestone #25 (Miner Wave 4 — AMS Hardening & Packaging): #5184 (Grafana datasource for AMS's local ledgers), #5185 (coding-agent usage/cost dashboard), and #5186-#5188 (Prometheus alert rules for AMS). A self-hoster running both ORB and AMS on one box should get ONE consolidated alerting pipeline, not two independently-built ones — keep this epic's design in mind when AMS's alert rules land.