eks-version-upgrade-readiness: add EKS Version Rollback readiness - #84
Conversation
- new references/rollback-readiness.md: inverse per-version lookup bounded to the supported-version window, Auto Mode disruption blockers with insight severity, add-on cross-compatibility, rollbackConfig support per IaC tool - eks-specific-changes.md: add the missing 1.34 section, complete 1.35/1.36, order newest-first, add a Disruption Controls section - examples-before-after.md: 4 new examples (9-12) - SKILL.md: rollback window constraint, N-1 verdict per change, reference dispatch - README.md: Rollback Readiness section, restore the missing Known Limitations heading
…adiness, 1.33 to 1.34)
|
Benchmarks are in - Run 4: rollback readiness fixture, 1.33 -> 1.34The hop was picked on purpose. 1.33 -> 1.34 is the boundary where 18-file fixture: 13 manifests, a Terraform config (cluster + node group + Fargate profile), and a Helm chart. 14 planted cases spanning forward blockers, rollback blockers, flag-only, negative controls, and one deliberately out-of-range case. 14/14 behaved as specified, 11 automatic transformations, The result worth highlightingRather than take the skill's word for the N-1 verdicts, I ran the transformed manifests through The single extra skip on 1.33 is So the Second negative confirmation, for the flag-only rule: PodSecurityPolicy has no schema at 1.34 either ( The version-scoping control also held: the non-canonical CIDR case is a 1.36 issue, and for a 1.34 target the report files it as a future item explicitly marked "Not blocking for 1.34" rather than as a blocker. Three findings recorded honestly in BENCHMARKS.md
Validation commands were added to the |
…sight severity from this skill's own rating Two review follow-ups recorded in BENCHMARKS.md run 4. Add-on matrix: the single table mixed AWS-published per-version data with community floors, which is what led the agent to state a v2.8.0+ LB Controller minimum for 1.34 by extrapolating the 1.30+/1.32+ columns. Adding a 1.34+ column would have made it worse, since AWS publishes no per-Kubernetes-version floor for the LB Controller at all. Now split: an exact per-version table for kube-proxy, CoreDNS and VPC CNI cited to the AWS docs (1.31 through 1.36), and a floors table for self-managed add-ons with a source-of-truth link per row and an explicit instruction not to extrapolate. Severity: ERROR/WARNING are cluster insight severities and only apply to what EKS evaluates. A repository finding has no insight, so its Impact column is this skill's own risk rating. rollback-readiness.md now says which is which.
|
Closing the loop on the two follow-ups I flagged above: both are fixed in this PR ( On the add-on matrix specifically, I did not add the "1.34+ column" I offered earlier, because checking the source made it the wrong fix: AWS publishes no per-Kubernetes-version floor for the AWS Load Balancer Controller at any version, only a general "2.7.2 or later" recommendation. A 1.34+ column would have dressed up the same extrapolation as documented guidance. What the table does now is split by who publishes the number:
That targets the actual defect. The agent extrapolated because one table mixed two kinds of number, so the fix is structural rather than one more column. |
The hand-verified snapshot was measured against the live API one day after it was written: 17 of 18 cells had already drifted, and two of the changes were MINOR version bumps rather than build-suffix bumps (CoreDNS for k8s 1.33 v1.12.4 -> v1.13.2, VPC CNI v1.22.4 -> v1.23.0 on every row). Only kube-proxy for 1.36 was unchanged. Leads with aws eks describe-addon-versions as the authoritative source and demotes the table to a dated fallback with the measured drift stated, so a reader can tell the difference between a fact about their cluster and an artifact of this document. The two-section split (managed vs self-managed) and the do-not-extrapolate rule for self-managed floors are unchanged - they were already correct.
Summary
Amazon EKS Version Rollback (announced 2026-07-01) reverts the control plane one minor version within 7 days of an in-place upgrade. That changes what "upgrade-ready code" means: during the window, the deployed code has to be valid on both versions, so forward compatibility alone is no longer sufficient.
This PR adds rollback readiness to
eks-version-upgrade-readinessas a detection and report dimension. The skill still never touches a cluster — it does not callupdate-cluster-version,list-insights,describe-insightorcancel-update, and that boundary is now written into the Non-Goals.Why this belongs in a code-readiness skill
ROLLBACK_READINESSinsights only run after the upgrade, and only for 7 days. This skill runs before, so a blocker can be designed around instead of discovered during an incident.karpenter.sh/do-not-disrupt, PodDisruptionBudgets, node annotations. The same controls throttle the data plane upgrade, so the check pays off even with no rollback planned.What changed
references/rollback-readiness.mdrollbackConfigsupport per IaC tool, and an explicit statement of where severity is quoted from the EKS docs versus rated by the skill.references/eks-specific-changes.mdreferences/examples-before-after.mdapiVersionbump that closes the rollback window, and non-canonical IP/CIDR values.SKILL.mdREADME.md## Known Limitationsheading — the ToC links to#known-limitationsbut the heading was never in the file, so that anchor is currently a 404.BENCHMARKS.mdContent gaps fixed along the way
While cross-checking the release notes, three things turned out to be missing from
eks-specific-changes.mdindependently of rollback:gitRepovolume disablement,StrictIPCIDRValidationon by default, SELinux volume labeling GA, and theService.spec.externalIPsdeprecation.trafficDistribution: PreferSameNodeand StatefulSetmaxUnavailable.Maintenance: bounded on purpose
The inverse table is the one part of this that could grow without limit, so it is scoped to EKS versions in standard or extended support (currently 1.31 through 1.36). A rollback requires both sides of the hop to be supported, so a row outside that range describes something the API rejects anyway. When a version leaves extended support its row is deleted rather than kept, and the file carries the review date plus a link to the release calendar. This is deliberately the opposite of
api-removals-by-version.md, which stays historical from 1.16 because customers really do sit on very old manifests.Testing
nodes: "0"and node-level do-not-disrupt are ERROR; pod-level do-not-disrupt andmaxUnavailable: 0are WARNING) is taken from the Auto Mode rollback page, not inferred.rollbackConfigbounds (120 to 10080 minutes, default 720) confirmed against the EKS API reference andAWS::EKS::Cluster RollbackConfig.rollback_configargument exists inhashicorp/terraform-provider-aws(searchedinternal/service/eks, zero hits), so the file says report-only instead of showing HCL that would not plan. CDK exposes it on the L1CfnClusteronly.End-to-end run
BENCHMARKS.mdnow carries a fourth run measured against this content. Fixture: 18 files (13 manifests, Terraform including a Fargate profile, one Helm chart), upgrading 1.33 to 1.34. That hop was chosen deliberately — it is the boundary wherestorage.k8s.io/v1VolumeAttributesClass graduates, so a correct transformation is simultaneously a rollback blocker.14/14 planted cases behaved as specified, including the negative controls (6 files byte-identical against the git baseline), the flag-only rule (PodSecurityPolicy untouched), and one deliberately out-of-range case: a 1.36 IP/CIDR issue that must not be reported as blocking for a 1.34 target, and was reported as future and explicitly non-blocking.
The N-1 classification was confirmed empirically rather than asserted, by validating the same manifests against both versions' schemas:
That is the target-only verdict reproduced against upstream schemas with no cluster involved, which is why the validation commands now include running the schema check against both sides of the hop.
kubectl --dry-runwas not used in this run (the workstation kubeconfig pointed at a decommissioned cluster); run 1 already covers live-cluster acceptance against EKS 1.35. The same command independently reproduced that PodSecurityPolicy has no schema at 1.34, confirming the flag-only design from a second angle.Two follow-ups the run surfaced are fixed here rather than deferred:
v2.8.0+LB Controller minimum for 1.34, projected from the 1.30+ and 1.32+ columns. Adding a "1.34+" column would have made it worse, because AWS publishes no per-Kubernetes-version floor for the LB Controller at all — only a general "2.7.2 or later" recommendation. The table is now split by who publishes the number, and the self-managed half says explicitly not to extrapolate to a version with no column: report the floor as undocumented and cite the upstream source instead.ERRORto a manifest finding. No insight exists for a repository finding, so that label sends a reader hunting for something EKS will never surface.rollback-readiness.mdnow separates quoted severity from the skill's own risk rating.Budget note for maintainers: the reference set roughly doubled with the rollback material, so this transformation wants 90-120 agent minutes rather than 60. Run 4 measured 61.2 and was cut at its own final
terraform validateunder a 60-minute cap, with all deliverables already complete on disk.By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.