Skip to content

eks-version-upgrade-readiness: add EKS Version Rollback readiness - #84

Merged
venuvasu merged 4 commits into
aws-samples:mainfrom
brunokktro:feat/eks-rollback-readiness
Aug 31, 2026
Merged

venuvasu merged 4 commits into
aws-samples:mainfrom
brunokktro:feat/eks-rollback-readiness

Conversation

@brunokktro

@brunokktro brunokktro commented Aug 18, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Amazon EKS Version Rollback (announced 2026-07-01) reverts the control plane one minor version within 7 days of an in-place upgrade. That changes what "upgrade-ready code" means: during the window, the deployed code has to be valid on both versions, so forward compatibility alone is no longer sufficient.

This PR adds rollback readiness to eks-version-upgrade-readiness as a detection and report dimension. The skill still never touches a cluster — it does not call update-cluster-version, list-insights, describe-insight or cancel-update, and that boundary is now written into the Non-Goals.

Why this belongs in a code-readiness skill

Timing ROLLBACK_READINESS insights only run after the upgrade, and only for 7 days. This skill runs before, so a blocker can be designed around instead of discovered during an incident.
Coverage Per the cluster insights docs, the insights check EKS-managed add-ons only (CoreDNS, VPC CNI, kube-proxy), and not even those when the version was overridden outside the add-on lifecycle. Karpenter, the AWS Load Balancer Controller, Istio, Argo CD and cluster-autoscaler are the customer's responsibility, and this skill already reads their manifests and Helm values.
Surface Four of the Auto Mode checks are plain YAML the skill already parses: NodePool disruption budgets, karpenter.sh/do-not-disrupt, PodDisruptionBudgets, node annotations. The same controls throttle the data plane upgrade, so the check pays off even with no rollback planned.

What changed

File Change
references/rollback-readiness.md New. Inverse per-version lookup (what was added in N that does not exist in N-1), classified as resource / field / enum / gated / behavior with a rollback impact per row. Plus Auto Mode disruption blockers with their mapped insight severity, add-on cross-compatibility strategy, the staged CP-bake-DP sequence, rollbackConfig support per IaC tool, and an explicit statement of where severity is quoted from the EKS docs versus rated by the skill.
references/eks-specific-changes.md Adds the missing 1.34 section (it was absent entirely), completes 1.35 and 1.36, reorders newest-first (1.35/1.36 were appended after 1.28), and adds a Disruption Controls section. The add-on compatibility matrix is now split by who publishes the version: an exact per-Kubernetes-version table for the three EKS-managed add-ons, and a floors table for self-managed add-ons with a source-of-truth link per row. No existing bullet was removed.
references/examples-before-after.md 4 new examples (9-12): NodePool budget that blocks rollback, PDB that stalls node replacement, a correct apiVersion bump that closes the rollback window, and non-canonical IP/CIDR values.
SKILL.md New "Rollback Window Awareness" constraint, N-1 verdict per transformation, rollback phase in the workflow, reference dispatch row, two exit criteria, updated Non-Goals.
README.md New Rollback Readiness section, new limitations, updated design decisions. Also restores the missing ## Known Limitations heading — the ToC links to #known-limitations but the heading was never in the file, so that anchor is currently a 404.
BENCHMARKS.md Adds run 4, measured against this content (see below), plus two new exit criteria rows.

Content gaps fixed along the way

While cross-checking the release notes, three things turned out to be missing from eks-specific-changes.md independently of rollback:

  • No 1.34 section at all — so AL2 AMIs not being released for 1.34, containerd 2.1, VolumeAttributesClass GA, DRA GA and the cgroup-driver deprecation were undocumented.
  • 1.36 was incomplete — missing gitRepo volume disablement, StrictIPCIDRValidation on by default, SELinux volume labeling GA, and the Service.spec.externalIPs deprecation.
  • 1.35 was incomplete — missing trafficDistribution: PreferSameNode and StatefulSet maxUnavailable.

Maintenance: bounded on purpose

The inverse table is the one part of this that could grow without limit, so it is scoped to EKS versions in standard or extended support (currently 1.31 through 1.36). A rollback requires both sides of the hop to be supported, so a row outside that range describes something the API rejects anyway. When a version leaves extended support its row is deleted rather than kept, and the file carries the review date plus a link to the release calendar. This is deliberately the opposite of api-removals-by-version.md, which stays historical from 1.16 because customers really do sit on very old manifests.

Testing

  • Every version-specific claim was cross-checked against the AWS docs rather than release-note memory: the standard-support and extended-support release notes, the release calendar, cluster insights, and Auto Mode rollback.
  • Severity mapping (nodes: "0" and node-level do-not-disrupt are ERROR; pod-level do-not-disrupt and maxUnavailable: 0 are WARNING) is taken from the Auto Mode rollback page, not inferred.
  • rollbackConfig bounds (120 to 10080 minutes, default 720) confirmed against the EKS API reference and AWS::EKS::Cluster RollbackConfig.
  • Terraform: no rollback_config argument exists in hashicorp/terraform-provider-aws (searched internal/service/eks, zero hits), so the file says report-only instead of showing HCL that would not plan. CDK exposes it on the L1 CfnCluster only.
  • Add-on versions in the managed table are quoted from the AWS docs pages for kube-proxy, CoreDNS and VPC CNI, verified 2026-08-20.
  • Markdown: all fenced blocks balanced and language-tagged; example numbering is sequential 1-12; no existing content dropped (verified by diffing the pre-change bullet inventory).

End-to-end run

BENCHMARKS.md now carries a fourth run measured against this content. Fixture: 18 files (13 manifests, Terraform including a Fargate profile, one Helm chart), upgrading 1.33 to 1.34. That hop was chosen deliberately — it is the boundary where storage.k8s.io/v1 VolumeAttributesClass graduates, so a correct transformation is simultaneously a rollback blocker.

14/14 planted cases behaved as specified, including the negative controls (6 files byte-identical against the git baseline), the flag-only rule (PodSecurityPolicy untouched), and one deliberately out-of-range case: a 1.36 IP/CIDR issue that must not be reported as blocking for a 1.34 target, and was reported as future and explicitly non-blocking.

The N-1 classification was confirmed empirically rather than asserted, by validating the same manifests against both versions' schemas:

kubeconform -kubernetes-version 1.34.0  ->  Valid 11, Invalid 0, Skipped 2
kubeconform -kubernetes-version 1.33.0  ->  Valid 10, Invalid 0, Skipped 3

The one extra skip on 1.33 is the VolumeAttributesClass: valid at 1.34,
"could not find schema" at 1.33. Every other transformed resource validates
on BOTH versions, matching its "safe on both" verdict.

That is the target-only verdict reproduced against upstream schemas with no cluster involved, which is why the validation commands now include running the schema check against both sides of the hop. kubectl --dry-run was not used in this run (the workstation kubeconfig pointed at a decommissioned cluster); run 1 already covers live-cluster acceptance against EKS 1.35. The same command independently reproduced that PodSecurityPolicy has no schema at 1.34, confirming the flag-only design from a second angle.

Two follow-ups the run surfaced are fixed here rather than deferred:

  • The add-on matrix invited extrapolation. The generated report stated a v2.8.0+ LB Controller minimum for 1.34, projected from the 1.30+ and 1.32+ columns. Adding a "1.34+" column would have made it worse, because AWS publishes no per-Kubernetes-version floor for the LB Controller at all — only a general "2.7.2 or later" recommendation. The table is now split by who publishes the number, and the self-managed half says explicitly not to extrapolate to a version with no column: report the floor as undocumented and cite the upstream source instead.
  • Severity was conflated. The report attached insight severity ERROR to a manifest finding. No insight exists for a repository finding, so that label sends a reader hunting for something EKS will never surface. rollback-readiness.md now separates quoted severity from the skill's own risk rating.

Budget note for maintainers: the reference set roughly doubled with the rollback material, so this transformation wants 90-120 agent minutes rather than 60. Run 4 measured 61.2 and was cut at its own final terraform validate under a 60-minute cap, with all deliverables already complete on disk.

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

- new references/rollback-readiness.md: inverse per-version lookup bounded to
  the supported-version window, Auto Mode disruption blockers with insight
  severity, add-on cross-compatibility, rollbackConfig support per IaC tool
- eks-specific-changes.md: add the missing 1.34 section, complete 1.35/1.36,
  order newest-first, add a Disruption Controls section
- examples-before-after.md: 4 new examples (9-12)
- SKILL.md: rollback window constraint, N-1 verdict per change, reference dispatch
- README.md: Rollback Readiness section, restore the missing Known Limitations heading
@brunokktro

Copy link
Copy Markdown
Contributor Author

Benchmarks are in - BENCHMARKS.md now carries a 4th run, so the "not re-run end-to-end" caveat in the description above no longer applies.

Run 4: rollback readiness fixture, 1.33 -> 1.34

The hop was picked on purpose. 1.33 -> 1.34 is the boundary where storage.k8s.io/v1 VolumeAttributesClass graduates, so a correct transformation is simultaneously a rollback blocker - exactly the case this PR is about.

18-file fixture: 13 manifests, a Terraform config (cluster + node group + Fargate profile), and a Helm chart. 14 planted cases spanning forward blockers, rollback blockers, flag-only, negative controls, and one deliberately out-of-range case. 14/14 behaved as specified, 11 automatic transformations, MIGRATION_REPORT.md generated with the Rollback Readiness section. 61.2 agent minutes, $2.14.

The result worth highlighting

Rather than take the skill's word for the N-1 verdicts, I ran the transformed manifests through kubeconform against both sides of the hop:

kubeconform -kubernetes-version 1.34.0  ->  Valid 11, Invalid 0, Skipped 2
kubeconform -kubernetes-version 1.33.0  ->  Valid 10, Invalid 0, Skipped 3

The single extra skip on 1.33 is volumeattributesclass.yaml. Isolated:

1.34.0  rc=0  VolumeAttributesClass gp3-fast is valid
1.33.0  rc=1  failed validation: could not find schema for VolumeAttributesClass

So the target-only (closes the rollback window) verdict is reproduced against upstream schemas instead of inferred from release notes, and every other transformed resource validates on both versions, matching its safe on both verdict. Two resources are skipped rather than passed (the Karpenter NodePool CRD and PSP have no upstream schema) and are reported as skipped.

Second negative confirmation, for the flag-only rule: PodSecurityPolicy has no schema at 1.34 either (could not find schema for PodSecurityPolicy), independently reproducing what run 1 showed against a live cluster.

The version-scoping control also held: the non-canonical CIDR case is a 1.36 issue, and for a 1.34 target the report files it as a future item explicitly marked "Not blocking for 1.34" rather than as a blocker.

Three findings recorded honestly in BENCHMARKS.md

  1. Budget. The run needed 61.23 agent minutes against a 60 limit and was cut at the agent's own final terraform validate. Transformations and the report were already complete on disk but left uncommitted, so the audit ran against the working tree. The reference set roughly doubled with the rollback material - 90-120 agent minutes is the realistic budget now, not 60.
  2. Add-on matrix gap. The generated report states a v2.8.0+ minimum for the AWS Load Balancer Controller on 1.34, but the matrix only has 1.30+ and 1.32+ columns, so that number is an extrapolation. An explicit 1.34+ column would close it. Happy to add it here or in a follow-up, whichever you prefer.
  3. Wording nit. The report labels the VolumeAttributesClass adoption with insight severity ERROR. It is an API-compatibility finding that would surface as an ERROR insight, but the table conflates the skill's own classification with the cluster insight severity.

kubectl --dry-run was not used in this run (the workstation kubeconfig pointed at a decommissioned cluster); kubeconform against both schema sets covers the same ground without a cluster, and run 1 already covers live-cluster acceptance.

Validation commands were added to the Validation Commands Used section so the dual-version check is reproducible.

…sight severity from this skill's own rating

Two review follow-ups recorded in BENCHMARKS.md run 4.

Add-on matrix: the single table mixed AWS-published per-version data with community floors, which is what led the agent to state a v2.8.0+ LB Controller minimum for 1.34 by extrapolating the 1.30+/1.32+ columns. Adding a 1.34+ column would have made it worse, since AWS publishes no per-Kubernetes-version floor for the LB Controller at all. Now split: an exact per-version table for kube-proxy, CoreDNS and VPC CNI cited to the AWS docs (1.31 through 1.36), and a floors table for self-managed add-ons with a source-of-truth link per row and an explicit instruction not to extrapolate.

Severity: ERROR/WARNING are cluster insight severities and only apply to what EKS evaluates. A repository finding has no insight, so its Impact column is this skill's own risk rating. rollback-readiness.md now says which is which.
@brunokktro

Copy link
Copy Markdown
Contributor Author

Closing the loop on the two follow-ups I flagged above: both are fixed in this PR (cb76804), and the description now reflects the run-4 benchmark instead of the pre-run wording it still carried.

On the add-on matrix specifically, I did not add the "1.34+ column" I offered earlier, because checking the source made it the wrong fix: AWS publishes no per-Kubernetes-version floor for the AWS Load Balancer Controller at any version, only a general "2.7.2 or later" recommendation. A 1.34+ column would have dressed up the same extrapolation as documented guidance.

What the table does now is split by who publishes the number:

  • EKS-managed add-ons (kube-proxy, CoreDNS, VPC CNI) get an exact per-version table for 1.31 through 1.36, quoted from the AWS docs, with a note that these are the latest published builds rather than minimum floors.
  • Self-managed add-ons keep the range-based floors but each row now carries the upstream source of truth, plus an explicit instruction not to extrapolate to a version with no column: report the floor as undocumented and cite the upstream matrix instead.

That targets the actual defect. The agent extrapolated because one table mixed two kinds of number, so the fix is structural rather than one more column.

The hand-verified snapshot was measured against the live API one day after it was
written: 17 of 18 cells had already drifted, and two of the changes were MINOR version
bumps rather than build-suffix bumps (CoreDNS for k8s 1.33 v1.12.4 -> v1.13.2, VPC CNI
v1.22.4 -> v1.23.0 on every row). Only kube-proxy for 1.36 was unchanged.

Leads with aws eks describe-addon-versions as the authoritative source and demotes the
table to a dated fallback with the measured drift stated, so a reader can tell the
difference between a fact about their cluster and an artifact of this document.

The two-section split (managed vs self-managed) and the do-not-extrapolate rule for
self-managed floors are unchanged - they were already correct.
@venuvasu
venuvasu merged commit 2b06c0a into aws-samples:main Aug 31, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants