fix(clusterapi): bind local API EKS lifecycle actions to the creating region - #6385
Conversation
… region Delete, start and stop rebuilt the EKS target from the AWS region selected at action time, so changing the region in Settings after creating a cluster could redirect a destructive action at a same-named cluster in another region and overwrite the only local evidence of the original target. Post-create actions now resolve the region from the binding written at create, and refuse when that evidence is missing or inconsistent. Part of #6203
…e EKS cleanup test
✅MegaLinter analysis: Success✅ Linters with no issuesactionlint, bash-exec, git_diff, hadolint, jscpd, jsonlint, lychee, markdown-table-formatter, markdownlint, prettier, prettier, shellcheck, shfmt, stylelint, syft, trivy-sbom, trufflehog, v8r, v8r, yamllint Notices📣 MegaLinter 9.5.0 is out! Discover the new features and security recommendations in the release announcement. (Skip this info by defining See detailed reports in MegaLinter artifacts
|
Verification recordTraced the enacting path, because the first shape of this fix was a measured no-op. My initial change bound the region into the The region now reaches the AWS call, confirmed link by link on
Exercised as its user, not through fixtures. The tests drive the real Ablations — 3, each proven RED, then restored and re-verified GREEN (restore confirmed by content, not assumption):
Control that must NOT move: under A1 the create-path test ( Local checks: One existing test was changed, deliberately. |
Unparked — the recorded blocker was stale. This PR was parked as blocked on The advisory applied to the base, not this diff — this branch was simply behind No action is needed on #6341, which remains automation-owned and is now a no-op against current |
CI is complete and green at this head (42 checks, 0 failures), including @coderabbitai review |
|
✅ Action performedReview finished.
|
📝 WalkthroughWalkthroughEKS configuration generation now uses persisted ownership state and recorded region metadata for previously created clusters, while retaining ambient-region behavior before creation completes. Configuration paths are canonicalized under the local cluster directory, with directory creation limited to writes. Local lifecycle jobs reject overlapping operations and bind EKS mutations to persisted cluster targets, refusing missing or mismatched state. Tests cover region preservation, fail-closed behavior, mutation guards, and cleanup failures. Possibly related issues
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 5
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
pkg/cli/clusterapi/distconfig.go (1)
76-85: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winEmpty ambient
AWS_REGIONat create time bakes in an unrecoverable binding.If
AWS_REGIONis unset,writeEKSConfigstamps an emptymetadata.region. Once persisted state exists,boundEKSConfigrejects that file permanently ("records no region"), so every later delete/start/stop is refused and the only escape is the rebind command. Failing fast on create withapi.ErrInvalidis cheaper than a cluster that can never be deleted through KSail.🛡️ Proposed guard
region := os.Getenv(credentials.DefaultEnvVar(credentials.AWSRegion)) + if region == "" { + return nil, fmt.Errorf( + "%w: no AWS region selected for EKS cluster %q; set AWS_REGION before creating it", + api.ErrInvalid, name, + ) + } configPath, err := writeEKSConfig(name, region)🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@pkg/cli/clusterapi/distconfig.go` around lines 76 - 85, Validate the region read in the create flow before calling writeEKSConfig: when AWS_REGION resolves to an empty value, return api.ErrInvalid instead of persisting an EKS configuration with no region. Keep the existing writeEKSConfig and DistributionConfig construction unchanged for valid regions.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@pkg/cli/clusterapi/distconfig.go`:
- Around line 127-134: In the config-reading flow containing the os.ReadFile
call, replace it with the repository-standard fsutil.ReadFileSafe helper and
remove the gosec suppression comment. Preserve the existing error wrapping with
api.ErrInvalid, the cluster name, and the underlying read error.
- Around line 146-152: Update the error guidance in the parsed.Metadata.Region
validation path to reference a supported recovery workflow. Either register the
`cluster rebind-eks-ownership` command and ensure it re-establishes ownership
metadata, or replace the command text in the `fmt.Errorf` message with an
existing CLI recovery path; keep the invalid-configuration error behavior
unchanged.
In `@pkg/cli/clusterapi/local_service_test.go`:
- Around line 1314-1346: The test
TestEKSActionAfterCreateBindsRegionToCreationNotCurrentSettings only validates
ExportEKSConfigForCreate, not successful lifecycle actions. Add a positive
delete, start, or stop case that uses the persisted ownership state, changes
AWS_REGION afterward, invokes the corresponding action, and asserts the
provisioner receives the cluster’s creation region rather than the current
setting.
- Around line 1387-1454: Add a regression test covering service.Delete while the
same cluster’s Create job is still in progress, exercising startJob’s
jobInProgress check before ownership-state validation. Assert the operation
returns the “already in progress” error and does not report the
missing-ownership error or invoke the provisioner’s delete path, using the
existing test helpers and fake provisioner patterns.
In `@pkg/cli/clusterapi/local_service.go`:
- Around line 576-603: Update eksMutationTarget to use the bound on-disk EKS
configuration to resolve and validate the runtime region, while retaining
persisted ClusterSpec only for ownership checks. Build and return the action
target from the resolved EKS binding and registry data so delete/start/stop
receive the current runtime configuration rather than the persisted ownership
baseline.
---
Outside diff comments:
In `@pkg/cli/clusterapi/distconfig.go`:
- Around line 76-85: Validate the region read in the create flow before calling
writeEKSConfig: when AWS_REGION resolves to an empty value, return
api.ErrInvalid instead of persisting an EKS configuration with no region. Keep
the existing writeEKSConfig and DistributionConfig construction unchanged for
valid regions.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository YAML (base), Organization UI (inherited)
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 985fce0c-1ad9-4f85-beda-a41880a4e572
📒 Files selected for processing (3)
pkg/cli/clusterapi/distconfig.gopkg/cli/clusterapi/local_service.gopkg/cli/clusterapi/local_service_test.go
📜 Review details
🧰 Additional context used
📓 Path-based instructions (4)
**/*.go
📄 CodeRabbit inference engine (AGENTS.md)
**/*.go: Use Go 1.26.1 or newer, matching the version declared ingo.mod.
All user-supplied file path arguments in CLI commands must be canonicalized withfsutil.EvalCanonicalPathbefore use; create parent directories first for new output paths.
Usefsutil.ReadFileSafefor constrained file reads instead of reimplementing path-containment checks.
Do not manually register MCP or Copilot tool handlers; runnable Cobra commands are exposed through automatic generation inpkg/toolgen.
Use a typedexperimentalfield inksail.yamlfor configuration-gated behavior that is not an entire command; regenerate the schema and CRD.
Graduate validated experimental features by deleting the singleGuardcall; do not retain unnecessary experimental scaffolding.
Run formatting and linting withgolangci-lint run --fixandgolangci-lint run --timeout 5m; validate withgo buildandgo test ./....
Files:
pkg/cli/clusterapi/local_service.gopkg/cli/clusterapi/distconfig.gopkg/cli/clusterapi/local_service_test.go
pkg/cli/**/*.go
📄 CodeRabbit inference engine (AGENTS.md)
pkg/cli/**/*.go: New not-yet-stable commands must be wrapped withexperimental.Guard(cmd), remain disabled by default, and require the global--experimentalflag.
Test experimental commands in both states: enabled with--experimentaland disabled withexperimental.ErrDisabled.
Files:
pkg/cli/clusterapi/local_service.gopkg/cli/clusterapi/distconfig.gopkg/cli/clusterapi/local_service_test.go
**/*.{go,yaml,yml,md,mdx,ts,tsx,json}
📄 CodeRabbit inference engine (AGENTS.md)
Generated files must not be hand-edited; run
make generateas the canonical regeneration command.
Files:
pkg/cli/clusterapi/local_service.gopkg/cli/clusterapi/distconfig.gopkg/cli/clusterapi/local_service_test.go
**/*_test.go
📄 CodeRabbit inference engine (AGENTS.md)
Add regression tests for confident bug fixes and run flaky-test candidates repeatedly with
go test -run <T> -count=10 ./....
Files:
pkg/cli/clusterapi/local_service_test.go
🔇 Additional comments (5)
pkg/cli/clusterapi/distconfig.go (2)
188-201: LGTM!
122-125: 🎯 Functional CorrectnessNo change needed for the read path.
~/.ksail/clustersis created by the cluster workflow, andReadFileSafe/EvalCanonicalPathfailures would still be returned throughapi.ErrInvalidvia the call chain.> Likely an incorrect or invalid review comment.pkg/cli/clusterapi/local_service.go (2)
606-619: LGTM!
526-538: 🎯 Functional CorrectnessNo issue with
Createrouting throughstartJob.
Createhas its own local provisioned-cluster path and does not callstartJob.pkg/cli/clusterapi/local_service_test.go (1)
587-626: LGTM!
- refuse an EKS create when no AWS region is selected, instead of stamping an empty metadata.region that boundEKSConfig then rejects forever - read the bound eks.yaml through fsutil.ReadFileSafe, dropping a gosec suppression - point the recovery guidance at paths that exist (a command named in three messages was never registered) - keep the action target minimal: persisted state is the ownership baseline, not the runtime spec, and the region is carried by eks.yaml - cover the successful lifecycle path and the create-time region guard
Outside-diff finding (
So the name check runs first — extracted as Ablations: removing the region guard turns the refusal test RED; removing the name check ahead of it turns the precedence test RED. Both restored and the package is green. |
One verified finding on the new recovery guidance at
|
| command | message |
|---|---|
Start |
…delete it with the AWS tooling directly (eksctl delete cluster --name probe-eks --region <region>)… |
Stop |
…identical… |
Delete |
…identical… |
So an operator who asked to start a cluster is told to destroy it. That is the one direction
of this error that is not recoverable if followed.
The in-product recovery path still exists and is not mentioned
ksail cluster eks-bind re-establishes exactly this binding, and pkg/svc/eksidentity/identity.go:261
already recommends it verbatim for the same situation:
run `ksail cluster eks-bind --name %s --provider AWS --experimental`
All three tokens matter — the command is experimental-gated, and its guard rejects any provider but
AWS — which is also why the previous text (ksail cluster rebind-eks-ownership --name …) could
never work: that name is not registered at all, and cobra answers unknown flag: --name. Replacing
it was right; pointing the operator out of KSail is what I'd reconsider.
Suggestion: keep the eksctl note as the last resort for a cluster that genuinely cannot be
rebound, and lead with eks-bind, phrased so it does not presume the operator wanted a delete.
I also found — by ablation, on my own copy — that asserting the binding by reading eks.yaml back
after changing AWS_REGION is vacuous: nothing rewrites that file on its own, so the assertion
passes even with boundEKSConfig stubbed to return nil. Re-resolving through
ExportEKSConfigForCreate is what discriminates a bound target from an ambient one. Worth checking
the new tests here against that, since the same shape is easy to land.
No action needed from me on this branch — it is yours; I have abandoned my copy.
…elper golangci-lint's wrapcheck flagged the unwrapped error returned from clusterapi.ExportEKSConfigForCreate in the test package's factory helper. Wrap it with %w so the cause stays inspectable via errors.Is/As.
CI is complete and green at @coderabbitai review |
|
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@pkg/cli/clusterapi/local_service.go`:
- Around line 577-585: Update the ErrStateNotFound error message in
confirmEKSOwnership to recommend restoring the missing binding with `ksail
cluster eks-bind --name ... --provider AWS --experimental` as the primary
recovery path, rather than advising `eksctl delete cluster`. Keep the message
applicable to delete, start, and stop operations, and retain region-confirmation
guidance only if still relevant.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository YAML (base), Organization UI (inherited)
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: c9a4a4ea-61cf-4988-b4f8-d00d92391cb5
📒 Files selected for processing (3)
pkg/cli/clusterapi/distconfig.gopkg/cli/clusterapi/local_service.gopkg/cli/clusterapi/local_service_test.go
📜 Review details
🧰 Additional context used
📓 Path-based instructions (4)
**/*.go
📄 CodeRabbit inference engine (AGENTS.md)
**/*.go: Use Go 1.26.1 or newer, matching the version declared ingo.mod.
All user-supplied file path arguments in CLI commands must be canonicalized withfsutil.EvalCanonicalPathbefore use; create parent directories first for new output paths.
Usefsutil.ReadFileSafefor constrained file reads instead of reimplementing path-containment checks.
Do not manually register MCP or Copilot tool handlers; runnable Cobra commands are exposed through automatic generation inpkg/toolgen.
Use a typedexperimentalfield inksail.yamlfor configuration-gated behavior that is not an entire command; regenerate the schema and CRD.
Graduate validated experimental features by deleting the singleGuardcall; do not retain unnecessary experimental scaffolding.
Run formatting and linting withgolangci-lint run --fixandgolangci-lint run --timeout 5m; validate withgo buildandgo test ./....
Files:
pkg/cli/clusterapi/local_service.gopkg/cli/clusterapi/distconfig.gopkg/cli/clusterapi/local_service_test.go
pkg/cli/**/*.go
📄 CodeRabbit inference engine (AGENTS.md)
pkg/cli/**/*.go: New not-yet-stable commands must be wrapped withexperimental.Guard(cmd), remain disabled by default, and require the global--experimentalflag.
Test experimental commands in both states: enabled with--experimentaland disabled withexperimental.ErrDisabled.
Files:
pkg/cli/clusterapi/local_service.gopkg/cli/clusterapi/distconfig.gopkg/cli/clusterapi/local_service_test.go
**/*.{go,yaml,yml,md,mdx,ts,tsx,json}
📄 CodeRabbit inference engine (AGENTS.md)
Generated files must not be hand-edited; run
make generateas the canonical regeneration command.
Files:
pkg/cli/clusterapi/local_service.gopkg/cli/clusterapi/distconfig.gopkg/cli/clusterapi/local_service_test.go
**/*_test.go
📄 CodeRabbit inference engine (AGENTS.md)
Add regression tests for confident bug fixes and run flaky-test candidates repeatedly with
go test -run <T> -count=10 ./....
Files:
pkg/cli/clusterapi/local_service_test.go
🔇 Additional comments (3)
pkg/cli/clusterapi/distconfig.go (1)
60-104: LGTM!Also applies to: 106-181, 183-218, 220-229, 238-267
pkg/cli/clusterapi/local_service.go (1)
509-554: LGTM!pkg/cli/clusterapi/local_service_test.go (1)
588-638: LGTM!Also applies to: 1016-1065, 1349-1385, 1426-1543
…sted action The ownership refusal is reached from every mutating entry point — Delete arrives as ClusterPhaseDeleting, Start and Stop both arrive as ClusterPhaseUpdating — but all three shared one message telling the operator to `eksctl delete cluster`. An operator who asked to start or stop a cluster was handed a destructive recovery step for a cluster KSail had just said it could not identify. Pass the phase into confirmEKSOwnership and select the guidance from it: deleting is suggested only on the delete path, while start/stop get non-destructive node-group guidance. The start/stop text names no KSail subcommand, because none exists to re-establish the binding.
Finding resolved at
|
One verified finding on the recovery guidance at
|
…bind command The refusal now names `ksail cluster eks-bind` first on every path. It restores the missing binding from the cluster the current AWS credentials select and never deletes or scales anything, so the refused action can simply be retried — which makes it the right first step for delete, start and stop alike. eksidentity.migrationRequiredError already points at the same command for the same condition, so the two now agree. Deletion stays offered only on the delete path. The test pins the exact command string on all three paths. The previous round on this PR shipped guidance naming a command the CLI does not expose, and my first attempt at this fix asserted no such command existed at all; both mistakes are now a test failure rather than something only a reviewer catches. Verified by running the binary: `ksail cluster eks-bind --experimental` resolves and accepts --name/--provider.
# Conflicts: # pkg/cli/clusterapi/local_service_test.go
Finding accepted — and it corrected an error of mine, at
|
Create is asynchronous, so asserting the retry succeeds without waiting for it left the goroutine running past the end of the test. Its state write then landed after t.Setenv and t.TempDir had unwound, creating .ksail/clusters under the real home directory — which is what the Home Isolation Guard caught. The wait is extracted as requireEventuallyPhase, since the same list-and-check block appears throughout this file and the reason it is required is worth stating once. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Readiness — what I exercised, as the operator rather than the authorThe change is only worth anything if the message actually reaches the person who has to go and look in {"error":"cleared a failed EKS create without deleting remote resources: \"failed-eks\" never recorded ownership state, so KSail could not identify the remote cluster. Check the AWS account for a partially created cluster of that name and delete it there. The failed job has been cleared, so create can be retried"}Measured in the same run, not read off the source:
Judged as its user: it names the cluster, says why KSail cannot delete it, gives the one action only a Both deltas, per the two-sided test. Floor: a Delete can no longer assert that a cluster is gone Validation: package green including |
Head @coderabbitai review |
|
Tip For best results, initiate chat on the files or code changes. 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
echo '== Current head and changed files =='
git rev-parse HEAD
git diff --name-status origin/main...HEAD
echo
echo '== Relevant outlines =='
for f in \
pkg/cli/clusterapi/local_service.go \
pkg/cli/clusterapi/local_service_test.go \
pkg/cli/clusterapi/export_test.go \
pkg/cli/clusterapi/distconfig.go
do
if [ -f "$f" ]; then
echo "--- $f ---"
ast-grep outline "$f" --items all --view expanded
fi
done
echo
echo '== Failed-create deletion, retry, and error-status references =='
rg -n -C 8 --glob '*.go' \
'clearedFailedEKSCreate|ErrEKSCreateClearedLocally|ClusterPhaseProvisioning|retry|Retry|Service\) Delete|func .*Delete|clientErrorStatus|writeError|SaveClusterSpec' \
pkg
echo
echo '== Focused current diff =='
git diff --no-ext-diff --unified=100 origin/main...HEAD -- \
pkg/cli/clusterapi/local_service.go \
pkg/cli/clusterapi/local_service_test.go \
pkg/cli/clusterapi/export_test.go \
pkg/cli/clusterapi/distconfig.goLength of output: 100062 🏁 Script executed: #!/bin/bash
set -euo pipefail
echo '== New failed-create delete path and completion paths =='
nl -ba pkg/cli/clusterapi/local_service.go | sed -n '380,455p;720,890p'
echo
echo '== New failed-create delete tests and shared wait helper =='
rg -n -C 6 --glob '*_test.go' \
'requireEventuallyPhase|TestDeleteEKSClearsFailedCreateWithoutOwnershipState|ErrEKSCreateClearedLocally|retried create|retry' \
pkg/cli/clusterapi
echo
echo '== HTTP error mapping definitions and tests =='
rg -n -C 12 --glob '*.go' \
'func clientErrorStatus|clientErrorStatus\(|func writeError|writeError\(' \
pkg
echo
echo '== Current-head commits and compact diff for the latest change =='
git log --oneline --decorate -8
git diff --no-ext-diff --stat HEAD~1..HEAD
git diff --no-ext-diff --unified=60 HEAD~1..HEAD -- \
pkg/cli/clusterapi/local_service.go \
pkg/cli/clusterapi/local_service_test.goLength of output: 1288 🏁 Script executed: #!/bin/bash
set -euo pipefail
numbered() {
awk -v start="$2" -v end="$3" \
'NR >= start && NR <= end { printf "%6d %s\n", NR, $0 }' "$1"
}
echo '== New failed-create delete path and completion paths =='
numbered pkg/cli/clusterapi/local_service.go 380 455
numbered pkg/cli/clusterapi/local_service.go 720 890
echo
echo '== New failed-create delete tests and shared wait helper =='
rg -n -C 8 --glob '*_test.go' \
'requireEventuallyPhase|TestDeleteEKSClearsFailedCreateWithoutOwnershipState|ErrEKSCreateClearedLocally|retried create|retry' \
pkg/cli/clusterapi
echo
echo '== HTTP error mapping definitions and tests =='
rg -n -C 12 --glob '*.go' \
'func clientErrorStatus|clientErrorStatus\(|func writeError|writeError\(' \
pkg
echo
echo '== Current-head commits and compact diff for the latest change =='
git log --oneline --decorate -8
git diff --no-ext-diff --stat HEAD~1..HEAD
git diff --no-ext-diff --unified=60 HEAD~1..HEAD -- \
pkg/cli/clusterapi/local_service.go \
pkg/cli/clusterapi/local_service_test.goLength of output: 61543
P1 — Persist ownership state for local API EKS creates
After a service restart or an Persist the immutable EKS ownership record after the local create succeeds. Include the recorded You are interacting with an AI system. |
CodeRabbit answered the 12:35Z request on its chat surface (an Analysis chain plus the State at this head: CI 25/25 green, 0 unresolved threads. The @codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 06ac198c96
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The region binding this PR adds is a lifecycle-only rule, but Create gates only on live discovery and in-flight jobs. A cluster removed out of band leaves spec.json and eks.yaml behind while discovery stops reporting it, so Create accepted the name again and the stale binding steered the new cluster into the region the old one was created in. Refuse instead of re-rendering: per-provider discovery failures are logged and skipped, so absence from the live set is not proof the remote cluster is gone, and re-rendering would overwrite the only local evidence binding a cluster that may still be running. The refusal names the delete command that clears the state, which succeeds even when the cluster is already gone.
Create had grown past the length limit with the stale-state guard. Extracting the reservation also removes the unlock/relock the guard first introduced: the state read needs no lock, only a defined position relative to the live check.
Head moved to That review's P1 was valid and is fixed: the region binding this PR adds was steering creates, not CI is green at this head and all 18 threads are resolved. @coderabbitai review |
|
Tip For best results, initiate chat on the files or code changes. 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
echo '== Requested/current head =='
git rev-parse HEAD
git status --short
git diff --name-status origin/main...HEAD
echo
echo '== Changed Go structure =='
git diff --name-only origin/main...HEAD -- '*.go' | while IFS= read -r f; do
echo "--- $f ---"
ast-grep outline "$f" --items all --view expanded || true
done
echo
echo '== Focused implementation and test references =='
rg -n -C 8 --glob '*.go' \
'reserve|reservation|startJob\(|runCreate\(|Create\(|resolveCluster\(|boundEKSConfig|bindFromOwnershipRecord|ListEKSOwnershipStates|SaveEKSOwnershipState|Capture\(|Persist\(|EKSOwnership|ClusterPhaseProvisioning|ClusterPhaseFailed' \
pkg/cli/clusterapi pkg/cli/cmd/cluster pkg/svc/state pkg/svc/eksidentity
echo
echo '== Current diff =='
git diff --no-ext-diff --unified=100 origin/main...HEAD -- \
pkg/cli/clusterapi \
pkg/svc/state \
pkg/svc/eksidentity \
pkg/cli/cmd/clusterLength of output: 99713 🏁 Script executed: #!/bin/bash
set -euo pipefail
echo '== Create and reservation implementation =='
rg -n -C 18 --glob '*.go' \
'func \(s \*Service\) Create\(|reserveCreateJob\(|staleStateErr|completed EKS|EKS create state|SaveClusterSpec\(' \
pkg/cli/clusterapi/local_service.go
echo
echo '== EKS config and ownership implementation =='
ast-grep outline pkg/cli/clusterapi/distconfig.go --items all --view expanded
rg -n -C 16 --glob '*.go' \
'func (eksDistributionConfig|boundEKSConfig|bindFromOwnershipRecord|eksCreateCompleted|writeEKSConfig)|ListEKSOwnershipStates\(|SaveEKSOwnershipState\(' \
pkg/cli/clusterapi/distconfig.go pkg/svc/state pkg/svc/eksidentity pkg/cli/cmd/cluster
echo
echo '== Current create-related tests =='
rg -n -C 12 --glob '*_test.go' \
'TestCreateRefusesANameWhoseEKSCreateStateRemains|ownership.*local|local.*ownership|new Service|different.*region|SaveEKSOwnershipState|ListEKSOwnershipStates|Service\.Create' \
pkg/cli/clusterapi
echo
echo '== Latest commits and latest diff =='
git log --oneline -6
git diff --no-ext-diff --unified=80 HEAD~2..HEAD -- pkg/cli/clusterapi/local_service.go pkg/cli/clusterapi/local_service_test.go pkg/cli/clusterapi/distconfig.goLength of output: 134348 🏁 Script executed: #!/bin/bash
set -euo pipefail
echo '== Lifecycle registration and factory timing =='
nl -ba pkg/cli/clusterapi/local_service.go | sed -n '520,655p;930,1035p'
rg -n -C 12 --glob '*.go' \
'func \(s \*Service\) (buildProvisioner|startJob|runDelete|runLifecycleAction)|distributionConfig\(|newFactory\(' \
pkg/cli/clusterapi
echo
echo '== Service-level stale/ambiguous lifecycle tests =='
rg -n -C 12 --glob '*_test.go' \
'Delete.*(stale|ambiguous|two-regions)|Service\.Delete|provisioner.*built|JobPresentForTest|ClusterPhaseDeleting|ClusterPhaseFailed' \
pkg/cli/clusterapi/local_service_test.go pkg/cli/clusterapi/export_test.go
echo
echo '== Local API EKS ownership persistence call sites =='
rg -n -C 8 --glob '*.go' \
'func \(s \*Service\) runCreate|eksidentity\.Capture|eksidentity\.Persist|SaveEKSOwnershipState\(' \
pkg/cli/clusterapi pkg/cli/cmd/cluster pkg/svc/eksidentityLength of output: 1288 🏁 Script executed: #!/bin/bash
set -euo pipefail
numbered() {
awk -v start="$2" -v end="$3" \
'NR >= start && NR <= end { printf "%6d %s\n", NR, $0 }' "$1"
}
echo '== Lifecycle registration and factory timing =='
numbered pkg/cli/clusterapi/local_service.go 520 655
numbered pkg/cli/clusterapi/local_service.go 930 1035
rg -n -C 12 --glob '*.go' \
'func \(s \*Service\) (buildProvisioner|startJob|runDelete|runLifecycleAction)|distributionConfig\(|newFactory\(' \
pkg/cli/clusterapi
echo
echo '== Service-level stale/ambiguous lifecycle tests =='
rg -n -C 12 --glob '*_test.go' \
'Delete.*(stale|ambiguous|two-regions)|Service\.Delete|provisioner.*built|JobPresentForTest|ClusterPhaseDeleting|ClusterPhaseFailed' \
pkg/cli/clusterapi/local_service_test.go pkg/cli/clusterapi/export_test.go
echo
echo '== Local API EKS ownership persistence call sites =='
rg -n -C 8 --glob '*.go' \
'func \(s \*Service\) runCreate|eksidentity\.Capture|eksidentity\.Persist|SaveEKSOwnershipState\(' \
pkg/cli/clusterapi pkg/cli/cmd/cluster pkg/svc/eksidentityLength of output: 93462
You are interacting with an AI system. |
CodeRabbit answered the request as chat rather than queueing a review — its reply carries State at this head: 26 checks pass, 0 failing, all 18 threads resolved, What to look at: @codex review |
Codex reacted to the trigger, so it accepted the request — but produced no artifact in 18 minutes. Worth recording that CodeRabbit answered this PR as chat twice today — 12:36:43Z at The next comment is a bare |
|
@cursor review |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_9951b51f-5f6a-4f84-b5e3-fac1312d4f80) |
devantler
left a comment
There was a problem hiding this comment.
🤖 Generated by the Agentic Engineer
Self-review (fallback — CodeRabbit, Codex and Cursor Bugbot unavailable)
Reviewed commit: 3e2f70b622847f36c16e7300c59e388d6b8f1371
All three external lanes were tried at this head and none delivered a usable review:
| Lane | Outcome |
|---|---|
| CodeRabbit | Answered as chat, twice today — 12:36:43Z at 06ac198c and 14:16:52Z at 3e2f70b6. Both carry > [!TIP] For best results, initiate chat on the files or code changes plus an analysis chain, and neither produced a review object at head. Real work, never queued as a review. |
| Codex | Reacted to the 14:24:44Z trigger, then produced nothing for 18 minutes. Its envelope on this PR is measurable — the 12:41:57Z request returned a P1 at 12:48:51Z, about 7 minutes — so this is a stall at ~2.5×, not normal latency. |
| Cursor Bugbot | conclusion: neutral + output.title: "Error" at 14:44:50Z: Bugbot couldn't run - usage limit reached. No retry window; only a Cursor admin can lift it. |
A usage/spend limit and a lane that answers as chat are both explicitly not-delivering conditions, so this local round substitutes for the bot review under the contract's fallback rule.
What I checked
Correctness and security of the whole PR at this head, not only my own commits: the create-refusal guard and its interaction with the region binding; the binding's precedence rules; the config-vs-ownership reconciliation; path containment in writeEKSConfig/canonicalClusterDir; and error-path behaviour under an unreadable state store.
Verdict: 1 finding (P0: 0, P1: 0, P2: 1)
The P2 is a pre-existing conflation in pkg/svc/state, which this PR does not touch but now depends on. Filed as #6429 rather than expanded into this PR's scope. It does not block promotion: without this PR there is no binding at all, so the change is strictly safer than main in every case — it is an incomplete guard, not a new hazard.
Verification performed
- RED/GREEN on the create refusal, then ablated in place at the final implementation shape (guard removed → the test fails on the deterministic job-table assertion, not on an async race).
- Full
pkg/cli/clusterapisuite green, including the existing lifecycle-binding controls that this PR exists for — so the refusal does not over-correct and break delete/start/stop binding. golangci-lint: zero new findings from this diff. The threegosecG704 reports reproduce identically at the base commit in files this PR never touches.- Two failures in
pkg/cli/cmd/open/chatunder the full./pkg/cli/...run were verified pre-existing by running the same command at the base commit06ac198c, where they reproduce identically. - Exercised as the operator: drove the real
Service.Createagainst real on-disk state with a stale binding and read the actual user-facing output —cluster already exists: "prod-eks" still has local state from a completed EKS create … Run \ksail cluster delete --name prod-eks` to remove the cluster and clear that state — it succeeds even when the cluster is already gone — then retry. The provisioner factory was never built, confirming the refusal precedes any provisioning. I also read the delete path to confirm that remedy is real:Deleteresolves the cluster from its ownership state andrunDeleteclears local state when the provisioner reportsErrClusterNotFound`.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3e2f70b622
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| completed, err := eksCreateCompleted(name) | ||
| if err != nil || !completed { | ||
| //nolint:nilerr // an undeterminable read is not proof of a completed create; see above. | ||
| return nil |
There was a problem hiding this comment.
Propagate conflicting distribution state before reserving the job
When a name has persisted Kind/Talos state but no live cluster, eksCreateCompleted returns the explicit distribution-conflict error, but this branch discards it and reserveCreateJob registers an EKS create. The background factory then encounters the same conflict and leaves a Failed job; clearedFailedEKSCreate cannot remove it because state exists, while every subsequent Create sees the job, so the web user cannot clear or retry the name. The fresh evidence relative to the earlier state-type finding is this newly added suppression: keep unreadable-state errors deferred if desired, but propagate the confirmed non-EKS name collision before registering the job.
Useful? React with 👍 / 👎.
| return fmt.Errorf( | ||
| "%w: %q still has local state from a completed EKS create, which binds it to the region it"+ | ||
| " was created in. Creating it again would provision new resources in that region rather"+ | ||
| " than the one selected now. Run `ksail cluster delete --name %s` to remove the cluster"+ | ||
| " and clear that state — it succeeds even when the cluster is already gone — then retry", |
There was a problem hiding this comment.
Surface state-cleanup failures before blocking recreation
When an EKS delete removes the AWS cluster but DeleteClusterState fails (for example because ~/.ksail/clusters becomes read-only), runDelete deliberately logs only a server-side warning and clears the job. This new guard then treats the surviving spec.json as a completed create and refuses every attempt to recreate the now-absent cluster; the UI no longer has a row to delete, and repeating deletion continues to report asynchronous success while leaving the blocking state. Either make the completed delete surface and retain the cleanup failure until state is actually removed, or provide a recovery path that clears this local-only residue.
Useful? React with 👍 / 👎.
| spec, err := state.LoadClusterSpec(name) | ||
| if err != nil { | ||
| if errors.Is(err, state.ErrStateNotFound) { | ||
| // Not created yet: the caller renders from the ambient region. | ||
| return false, nil |
There was a problem hiding this comment.
Treat immutable ownership records as completed creates
When spec.json is absent but a valid EKS ownership record remains—such as after ksail cluster eks-bind, or when the CLI successfully persisted immutable identity but its later best-effort SaveClusterSpec failed—this function reports that creation never completed. If Settings now selects another region, discovery misses the recorded cluster, the create guard allows the name, and boundEKSConfig also skips the ownership record and renders a fresh config, permitting a second same-named cluster to be provisioned in the new region. A usable immutable ownership record is positive evidence of a completed create and must cause Create to refuse the name even without spec.json.
Useful? React with 👍 / 👎.

Why
Creating an EKS cluster from the local web UI binds it to the AWS region selected in Settings at that moment. Deleting, starting or stopping it later did not reuse that binding — it rebuilt the target from whatever region was selected at action time.
So changing the region in Settings between creating a cluster and operating on it could point a destructive action at a different, same-named cluster in the newly selected region. The same code path also overwrote the only local record of the original target, so the evidence needed to notice the mistake was destroyed by the action itself.
What
Actions on an EKS cluster that has finished creating now resolve their target from the binding written when it was created, instead of from the current selection. When that evidence is missing or disagrees with the cluster being acted on, the action is refused rather than falling back to an unconfirmed target — falling back is exactly the redirect being prevented, and delete cannot be undone.
A refusal now tells the operator how to recover, and the advice matches what they asked for: it leads with the KSail command that restores the missing binding — which touches nothing in AWS, so the original action can simply be retried — and only suggests deleting the cluster when deleting is what was requested. Previously every refusal, including start and stop, advised deleting a cluster KSail had just said it could not identify.
Creating a cluster is unchanged, including retrying after a failed create: until a create succeeds there is no binding, so the region selected now still applies and a corrected region is never ignored.
Part of #6203 — the first slice. The remaining acceptance criteria (an exact read-only ownership query before each mutation, and alignment with the immutable AWS account/cluster identity in #6202) need AWS identity calls and are deliberately not in this change.