Skill eval infra error — sre-lead (holdout)
Outcome: error — the eval could not score the skill (infrastructure failure: e.g. every model in the fallback chain throttled, or bad usage / missing tooling).
This is an infra signal, not a skill regression: the model never produced a scorable answer, so the held-out set is currently un-scored and needs a look. Report-only (Phase 1): it does not block any merge.
| skill |
score |
passed |
failed |
total |
sre-lead |
0 |
0 |
4 |
4 |
Regressed cases
sre-lead-hold-critical-path-no-slo — expected {"risk_tier":"HIGH","escalate":true,"recommend":"define an availability/latency SLO with an error budget, a severity tier, and a runbook before this critical-path endpoint launches"}, got null
sre-lead-hold-region-failover-no-dr — expected {"risk_tier":"HIGH","escalate":true,"recommend":"restore or re-plan a failover path with explicit RTO/RPO targets and a practiced failover test before removing the standby; quantify the single-region blast radius"}, got null
sre-lead-hold-blameful-postmortem — expected {"risk_tier":"MEDIUM","escalate":false,"recommend":"rewrite blamelessly around a timeline, root cause, and contributing factors, and add action items with named owners and deadlines"}, got null
sre-lead-hold-prod-chaos-no-safety — expected {"risk_tier":"HIGH","escalate":true,"recommend":"gate the experiment behind a steady-state hypothesis, a minimized blast radius (single service/off-peak), and an automatic abort before running any prod chaos"}, got null
View the workflow run
Auto-managed by notify-eval-health.sh — this issue is updated in place each run and closed automatically on recovery.
Skill eval infra error —
sre-lead(holdout)Outcome:
error— the eval could not score the skill (infrastructure failure: e.g. every model in the fallback chain throttled, or bad usage / missing tooling).This is an infra signal, not a skill regression: the model never produced a scorable answer, so the held-out set is currently un-scored and needs a look. Report-only (Phase 1): it does not block any merge.
sre-leadRegressed cases
sre-lead-hold-critical-path-no-slo— expected {"risk_tier":"HIGH","escalate":true,"recommend":"define an availability/latency SLO with an error budget, a severity tier, and a runbook before this critical-path endpoint launches"}, got nullsre-lead-hold-region-failover-no-dr— expected {"risk_tier":"HIGH","escalate":true,"recommend":"restore or re-plan a failover path with explicit RTO/RPO targets and a practiced failover test before removing the standby; quantify the single-region blast radius"}, got nullsre-lead-hold-blameful-postmortem— expected {"risk_tier":"MEDIUM","escalate":false,"recommend":"rewrite blamelessly around a timeline, root cause, and contributing factors, and add action items with named owners and deadlines"}, got nullsre-lead-hold-prod-chaos-no-safety— expected {"risk_tier":"HIGH","escalate":true,"recommend":"gate the experiment behind a steady-state hypothesis, a minimized blast radius (single service/off-peak), and an automatic abort before running any prod chaos"}, got nullView the workflow run
Auto-managed by
notify-eval-health.sh— this issue is updated in place each run and closed automatically on recovery.