Skip to content

OCPBUGS-86494: fix(nodepool): preserve KubeVirt userdata Secrets during NodePool rollout - #8581

Merged
openshift-merge-bot[bot] merged 3 commits into
openshift:mainfrom
amasolov:fix/kubevirt-keep-old-userdata-during-rollout
Jul 27, 2026
Merged

openshift-merge-bot[bot] merged 3 commits into
openshift:mainfrom
amasolov:fix/kubevirt-keep-old-userdata-during-rollout

Conversation

@amasolov

@amasolov amasolov commented May 24, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it:

On KubeVirt, the bootstrap userdata Secret in the HCP namespace (user-data-<nodepool>-<hash>) is shared by all VMs in a given NodePool generation. During a rolling update, HyperShift's secretJanitor and Token.cleanupOutdated() eagerly delete the old userdata Secret as soon as a new config version is computed. Any VMs from the previous generation that are still running will then fail CAPK reconciliation because their referenced Secret no longer exists.

This PR extends the existing AWS platform guard (which already preserves old userdata Secrets for a different reason) to also cover KubeVirt. With this change, old userdata Secrets are retained until the rollout finishes and old VMs are cleaned up.

Which issue(s) this PR fixes:

Fixes premature garbage collection of shared userdata Secrets on the KubeVirt provider during NodePool rolling updates.

Special notes for your reviewer:

  • The fix mirrors the existing AWS guard pattern. AWS preserves old userdata because of a CAPA bug (Do not return error if secret does not exist kubernetes-sigs/cluster-api-provider-aws#3805). KubeVirt needs it because the Secret is architecturally shared across VMs in a generation.
  • Long term, a more targeted cleanup (e.g. checking whether any old-generation VMs still reference the Secret) would be preferable, but this minimal change is safe and consistent with the existing pattern.

Manual testing

Verified on a real KubeVirt HostedCluster (OCP 4.20 management cluster with CNV 4.20.15, guest release 4.22.3, 2-replica NodePool). Triggered a rolling update by changing VM compute cores. The old userdata Secret was preserved throughout the rollout, both old machines were deleted cleanly, and the rollout completed with no errors. Full test evidence: #8581 (comment)

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Made with Cursor

Summary by CodeRabbit

  • Bug Fixes

    • Improved secret retention handling for KubeVirt platforms, ensuring user-data secrets are properly preserved during node pool operations and cluster updates.
  • Tests

    • Extended test coverage to validate secret lifecycle management behavior for KubeVirt platforms.

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Pipeline controller notification
This repo is configured to use the pipeline controller. Second-stage tests will be triggered either automatically or after lgtm label is added, depending on the repository configuration. The pipeline controller will automatically detect which contexts are required and will utilize /test Prow commands to trigger the second stage.

For optional jobs, comment /test ? to see a list of all defined jobs. To trigger manually all jobs from second stage use /pipeline required command.

This repository is configured in: LGTM mode

@openshift-ci-robot openshift-ci-robot added the jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. label May 24, 2026
@openshift-ci openshift-ci Bot added the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label May 24, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@amasolov: This pull request explicitly references no jira issue.

Details

In response to this:

What this PR does / why we need it:

On KubeVirt, the bootstrap userdata Secret in the HCP namespace (user-data-<nodepool>-<hash>) is shared by all VMs in a given NodePool generation. During a rolling update, HyperShift's secretJanitor and Token.cleanupOutdated() eagerly delete the old userdata Secret as soon as a new config version is computed. Any VMs from the previous generation that are still running will then fail CAPK reconciliation because their referenced Secret no longer exists.

This PR extends the existing AWS platform guard (which already preserves old userdata Secrets for a different reason) to also cover KubeVirt. With this change, old userdata Secrets are retained until the rollout finishes and old VMs are cleaned up.

Which issue(s) this PR fixes:

Fixes premature garbage collection of shared userdata Secrets on the KubeVirt provider during NodePool rolling updates.

Special notes for your reviewer:

  • The fix mirrors the existing AWS guard pattern. AWS preserves old userdata because of a CAPA bug (Do not return error if secret does not exist kubernetes-sigs/cluster-api-provider-aws#3805). KubeVirt needs it because the Secret is architecturally shared across VMs in a generation.
  • Long term, a more targeted cleanup (e.g. checking whether any old-generation VMs still reference the Secret) would be preferable, but this minimal change is safe and consistent with the existing pattern.

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Made with Cursor

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@coderabbitai

coderabbitai Bot commented May 24, 2026 •

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

This PR extends the existing user-data Secret retention logic to support KubeVirt platform NodePools in addition to AWS. The change adds an unconditional early return in shouldKeepOldUserData when the HostedCluster platform is KubevirtPlatform, followed by a parallel update to the cleanupOutdated function in token management to skip deletion for both AWS and KubeVirt. Test cases were added for both modules to verify that old user-data Secrets are preserved for KubeVirt platforms while token secrets receive expiration timestamps, and assertions were updated to conditionally expect retention or deletion based on platform type.

🚥 Pre-merge checks | ✅ 11 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (11 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Stable And Deterministic Test Names ✅ Passed Both new test names in secret_janitor_test.go and token_test.go are stable and deterministic, containing only static descriptive strings with no generated identifiers, timestamps, or dynamic values.
Test Structure And Quality ✅ Passed Tests use table-driven patterns with meaningful assertion messages and follow codebase conventions for Gomega/testing.T tests.
Microshift Test Compatibility ✅ Passed No Ginkgo e2e tests added in this PR. Changes only include standard Go unit tests (testing.T pattern) for controller logic changes to KubeVirt userdata Secret retention.
Single Node Openshift (Sno) Test Compatibility ✅ Passed PR adds standard Go unit tests with testing.T, not Ginkgo e2e tests. Check applies only to Ginkgo (It, Describe, Context), so it's not applicable.
Topology-Aware Scheduling Compatibility ✅ Passed PR modifies controller reconciliation logic for secret lifecycle, not scheduling constraints. No pod affinity, topology spread, nodeSelector, or scheduling specifications introduced.
Ote Binary Stdout Contract ✅ Passed OTE Binary Stdout Contract check is not applicable. Modified files are controller code and standard Go unit tests, not OTE/Ginkgo v2 test binaries. No stdout violations detected.
Ipv6 And Disconnected Network Test Compatibility ✅ Passed PR contains only standard Go unit tests (testing.T), not Ginkgo e2e tests. Custom check applies only to new Ginkgo e2e tests, therefore not applicable.
Title check ✅ Passed The title clearly and specifically describes the main change: preserving KubeVirt userdata Secrets during NodePool rollout, matching the core objective of the PR.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@openshift-ci openshift-ci Bot added area/hypershift-operator Indicates the PR includes changes for the hypershift operator and API - outside an OCP release needs-ok-to-test Indicates a PR that requires an org member to verify it is safe to test. labels May 24, 2026
@openshift-ci

openshift-ci Bot commented May 24, 2026

Copy link
Copy Markdown
Contributor

Hi @amasolov. Thanks for your PR.

I'm waiting for a openshift member to verify that this patch is reasonable to test. If it is, they should reply with /ok-to-test on its own line. Until that is done, I will not automatically test new commits in this PR, but the usual testing commands by org members will still work.

Regular contributors should join the org to skip this step.

Once the patch is verified, the new status will be reflected by the ok-to-test label.

I understand the commands that are listed here.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@hypershift-operator/controllers/nodepool/token.go`:
- Around line 181-187: The current guard unconditionally preserves KubeVirt
userdata Secrets; change it to preserve only while the old NodePool generation
is still referenced by existing VMs/Machines and delete once rollout completes.
Replace the permanent-platform check (t.nodePool.Spec.Platform.Type !=
hyperv1.KubevirtPlatform) with a rollout-aware condition that calls a helper
(e.g., add/use a function like isOldGenerationReferenced(nodePool) or
t.oldGenerationStillReferenced()) which inspects current Machine/VM objects or
NodePool status to detect references to the outdated generation, and only skip
deletion when that helper returns true; otherwise proceed to remove the
outdatedUserDataSecret() as normal.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 67ae5b4b-b88a-4532-a142-fe3cf9cabca4

📥 Commits

Reviewing files that changed from the base of the PR and between 6d994e4 and cd5f18b.

📒 Files selected for processing (4)
  • hypershift-operator/controllers/nodepool/secret_janitor.go
  • hypershift-operator/controllers/nodepool/secret_janitor_test.go
  • hypershift-operator/controllers/nodepool/token.go
  • hypershift-operator/controllers/nodepool/token_test.go

Comment on lines +181 to 187
// For AWS and KubeVirt, we keep the old userdata Secret so old Machines during rollout can be deleted.
// AWS: deletion fails because of https://github.com/kubernetes-sigs/cluster-api-provider-aws/pull/3805.
// KubeVirt: the Secret is shared by all VMs in the NodePool generation and must survive until
// the rollout completes and all old VMs are gone.
// TODO (Alberto): enable back deletion when the PR above gets merged.
if t.nodePool.Spec.Platform.Type != hyperv1.AWSPlatform {
if t.nodePool.Spec.Platform.Type != hyperv1.AWSPlatform && t.nodePool.Spec.Platform.Type != hyperv1.KubevirtPlatform {
userDataSecret := t.outdatedUserDataSecret()

@coderabbitai coderabbitai Bot May 24, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | 🏗️ Heavy lift

Make KubeVirt retention rollout-aware instead of permanent.

This guard keeps outdated KubeVirt userdata Secrets forever, not just during rollout. That diverges from the stated behavior (“preserve during rollout until old VMs are gone”) and can accumulate stale Secrets across upgrades. Please gate retention on an explicit “old generation still referenced” condition, then delete once rollout is complete.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@hypershift-operator/controllers/nodepool/token.go` around lines 181 - 187,
The current guard unconditionally preserves KubeVirt userdata Secrets; change it
to preserve only while the old NodePool generation is still referenced by
existing VMs/Machines and delete once rollout completes. Replace the
permanent-platform check (t.nodePool.Spec.Platform.Type !=
hyperv1.KubevirtPlatform) with a rollout-aware condition that calls a helper
(e.g., add/use a function like isOldGenerationReferenced(nodePool) or
t.oldGenerationStillReferenced()) which inspects current Machine/VM objects or
NodePool status to detect references to the outdated generation, and only skip
deletion when that helper returns true; otherwise proceed to remove the
outdatedUserDataSecret() as normal.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This intentionally follows the same unconditional platform guard that already exists for AWS. A rollout-aware check would be a larger refactor (both platforms would benefit) and is better suited as a follow-up.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@amasolov, fair point — keeping parity with the existing AWS unconditional guard is a reasonable approach for this minimal fix, and a rollout-aware cleanup that benefits both platforms is a sensible follow-up scope. Would you like me to open a GitHub issue to track that improvement (rollout-aware Secret cleanup for AWS and KubeVirt once old-generation VMs/Machines are gone)?


🧠 Learnings used
Learnt from: twolff-gh
Repo: openshift/hypershift PR: 8472
File: hypershift-operator/controllers/nodepool/azure.go:0-0
Timestamp: 2026-05-13T15:17:58.584Z
Learning: When constructing CAPZ `UserAssignedIdentity.ProviderID` values (used under `AzureMachineSpec.UserAssignedIdentities`), use the documented canonical `azure:///`-prefixed format: `azure:///subscriptions/{subscriptionId}/resourceGroups/{rg}/providers/Microsoft.ManagedIdentity/userAssignedIdentities/{name}`. CAPZ internally trims the prefix before calling the Azure SDK, so adding the prefix is still required by the contract. Use `capzutil.ProviderIDPrefix` ("azure:///") when building the ProviderID, and do not recommend removing or stripping this prefix in code review.

@amasolov amasolov changed the title NO-JIRA: fix(nodepool): preserve KubeVirt userdata Secrets during NodePool rollout OCPBUGS-86494: fix(nodepool): preserve KubeVirt userdata Secrets during NodePool rollout May 25, 2026
@amasolov
amasolov marked this pull request as ready for review May 25, 2026 23:56
@openshift-ci-robot openshift-ci-robot added jira/severity-important Referenced Jira bug's severity is important for the branch this PR is targeting. jira/invalid-bug Indicates that a referenced Jira bug is invalid for the branch this PR is targeting. labels May 25, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@amasolov: This pull request references Jira Issue OCPBUGS-86494, which is invalid:

  • expected the bug to target the "5.0.0" version, but no target version was set

Comment /jira refresh to re-evaluate validity if changes to the Jira bug are made, or edit the title of this pull request to link to a different bug.

The bug has been updated to refer to the pull request using the external bug tracker.

Details

In response to this:

What this PR does / why we need it:

On KubeVirt, the bootstrap userdata Secret in the HCP namespace (user-data-<nodepool>-<hash>) is shared by all VMs in a given NodePool generation. During a rolling update, HyperShift's secretJanitor and Token.cleanupOutdated() eagerly delete the old userdata Secret as soon as a new config version is computed. Any VMs from the previous generation that are still running will then fail CAPK reconciliation because their referenced Secret no longer exists.

This PR extends the existing AWS platform guard (which already preserves old userdata Secrets for a different reason) to also cover KubeVirt. With this change, old userdata Secrets are retained until the rollout finishes and old VMs are cleaned up.

Which issue(s) this PR fixes:

Fixes premature garbage collection of shared userdata Secrets on the KubeVirt provider during NodePool rolling updates.

Special notes for your reviewer:

  • The fix mirrors the existing AWS guard pattern. AWS preserves old userdata because of a CAPA bug (Do not return error if secret does not exist kubernetes-sigs/cluster-api-provider-aws#3805). KubeVirt needs it because the Secret is architecturally shared across VMs in a generation.
  • Long term, a more targeted cleanup (e.g. checking whether any old-generation VMs still reference the Secret) would be preferable, but this minimal change is safe and consistent with the existing pattern.

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Made with Cursor

Summary by CodeRabbit

  • Bug Fixes

  • Improved secret retention handling for KubeVirt platforms, ensuring user-data secrets are properly preserved during node pool operations and cluster updates.

  • Tests

  • Extended test coverage to validate secret lifecycle management behavior for KubeVirt platforms.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-ci openshift-ci Bot removed the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label May 25, 2026
@openshift-ci
openshift-ci Bot requested review from Nirshal and bryan-cox May 25, 2026 23:56
@amasolov

Copy link
Copy Markdown
Contributor Author

/jira refresh

@openshift-ci-robot openshift-ci-robot added jira/valid-bug Indicates that a referenced Jira bug is valid for the branch this PR is targeting. and removed jira/invalid-bug Indicates that a referenced Jira bug is invalid for the branch this PR is targeting. labels May 25, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@amasolov: This pull request references Jira Issue OCPBUGS-86494, which is valid. The bug has been moved to the POST state.

3 validation(s) were run on this bug
  • bug is open, matching expected state (open)
  • bug target version (5.0.0) matches configured target version for branch (5.0.0)
  • bug is in the state New, which is one of the valid states (NEW, ASSIGNED, POST)
Details

In response to this:

/jira refresh

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@bshirren

Copy link
Copy Markdown

/ok-to-test

@openshift-ci openshift-ci Bot added ok-to-test Indicates a non-member PR verified by an org member that is safe to test. and removed needs-ok-to-test Indicates a PR that requires an org member to verify it is safe to test. labels May 26, 2026
@codecov

codecov Bot commented May 26, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 44.51%. Comparing base (09265ac) to head (a2945ea).
⚠️ Report is 24 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #8581   +/-   ##
=======================================
  Coverage   44.50%   44.51%           
=======================================
  Files         774      774           
  Lines       96980    96986    +6     
=======================================
+ Hits        43164    43170    +6     
  Misses      50828    50828           
  Partials     2988     2988           
Files with missing lines Coverage Δ
...ft-operator/controllers/nodepool/secret_janitor.go 62.75% <100.00%> (+1.60%) ⬆️
hypershift-operator/controllers/nodepool/token.go 82.15% <100.00%> (ø)
Flag Coverage Δ
cmd-support 38.39% <ø> (ø)
cpo-hostedcontrolplane 47.19% <ø> (ø)
cpo-other 45.25% <ø> (ø)
hypershift-operator 54.45% <100.00%> (+0.01%) ⬆️
other 32.64% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

Rebase result: succeeded locally but push failed.

The rebase onto main completed cleanly (no conflicts), but the force-push to the fork failed with:

refusing to allow a GitHub App to create or update workflow `.github/workflows/docs-deploy.yaml` without `workflows` permission

The rebase pulled in upstream commits that modify .github/workflows/docs-deploy.yaml. The CI token does not have the workflows permission required to push workflow file changes to the fork.

To resolve this, the PR author can rebase locally and push:

git fetch upstream main
git rebase upstream/main
git push --force-with-lease origin fix/kubevirt-keep-old-userdata-during-rollout

Alternatively, enable "Allow edits from maintainers" on this PR and retry /rebase with a token that has the workflows permission.

amasolov and others added 3 commits July 23, 2026 09:46
…lout

On KubeVirt, the bootstrap userdata Secret in the HCP namespace is
shared by all VMs in a given NodePool generation. Premature deletion
of this Secret during a rolling update causes CAPK to fail
reconciliation for VMs that still reference it.

Extend the existing AWS platform guard in both shouldKeepOldUserData()
and cleanupOutdated() to also cover KubevirtPlatform, deferring
Secret cleanup until the rollout is complete.

Signed-off-by: Alexey Masolov <amasolov@redhat.com>
Assisted-by: Claude Opus 4.6 (via Cursor)
Co-authored-by: Cursor <cursoragent@cursor.com>
Address review feedback: distinguish the AWS temporary workaround
(CAPA bug, removable when OCP < 4.16 support is dropped) from the
KubeVirt architectural requirement in the code comments.

Signed-off-by: Alexey Masolov <amasolov@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…erData

Replace the if-chain with a switch statement and extract the AWS version
check into its own method for clarity. Also align the doc comments with
the style already used in token.go (distinguishing the AWS temporary
workaround from the KubeVirt architectural requirement).

Signed-off-by: Alexey Masolov <amasolov@redhat.com>
Assisted-by: Claude Opus 4.6 (via Cursor)
Co-authored-by: Cursor <cursoragent@cursor.com>
@amasolov
amasolov force-pushed the fix/kubevirt-keep-old-userdata-during-rollout branch from 6aba3d8 to a2945ea Compare July 22, 2026 23:47
@openshift-ci-robot openshift-ci-robot removed the verified Signifies that the PR passed pre-merge verification criteria label Jul 22, 2026
@openshift-ci openshift-ci Bot removed the lgtm Indicates that a PR is ready to be merged. label Jul 22, 2026
@amasolov

Copy link
Copy Markdown
Contributor Author

@bryan-cox I've rebased this branch onto the latest main. Not sure why the bot flagged "allow edits by maintainers" as it was already enabled on this PR.

@bshirren

Copy link
Copy Markdown

/ok-to-test

@bryan-cox

Copy link
Copy Markdown
Member

/lgtm

Putting back on from rebase

@openshift-ci openshift-ci Bot added the lgtm Indicates that a PR is ready to be merged. label Jul 23, 2026
@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Scheduling tests matching the pipeline_run_if_changed or not excluded by pipeline_skip_if_only_changed parameters:
/test e2e-aks-4-22
/test e2e-aws-4-22
/test e2e-aks
/test e2e-aws
/test e2e-aws-upgrade-hypershift-operator
/test e2e-azure-v2-self-managed
/test e2e-kubevirt-aws-ovn-reduced
/test e2e-v2-aws
/test e2e-v2-gke
/test unit
/test verify

@bryan-cox

Copy link
Copy Markdown
Member

/ok-to-test
/retest

@amasolov

Copy link
Copy Markdown
Contributor Author

@bryan-cox failing on the flaky tests again

@bryan-cox

Copy link
Copy Markdown
Member

@bryan-cox failing on the flaky tests again

@amasolov please take a look here and follow the workflow if you find the flaky tests are not related to your PR - https://hypershift.pages.dev/how-to/ci/triage/presubmit-failures/

@amasolov

Copy link
Copy Markdown
Contributor Author

@bryan-cox followed the triage workflow: https://hypershift.pages.dev/how-to/ci/triage/presubmit-failures/

Failing required jobs (commit a2945ea):

Relation to this PR: none. Diff is only KubeVirt userdata Secret retention in secret_janitor.go / token.go. e2e-kubevirt-aws-ovn-reduced passed.

Job history: both jobs are mostly red across other PRs right now (~14/20 failures each), so per the flowchart this looks like an infra/flake issue to escalate rather than a PR code change.

Will escalate in #forum-ocp-hypershift. Also noting verified dropped after the rebase.

@bshirren

Copy link
Copy Markdown

/retest-required

@bryan-cox

Copy link
Copy Markdown
Member

/test e2e-aws

@amasolov

Copy link
Copy Markdown
Contributor Author

/verified by @amasolov

@openshift-ci-robot

Copy link
Copy Markdown

@amasolov: Jira verification commands are restricted to collaborators for this repo.

Details

In response to this:

/verified by @amasolov

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@bryan-cox

Copy link
Copy Markdown
Member

/verified by @amasolov

See #8581 (comment)

@openshift-ci-robot openshift-ci-robot added the verified Signifies that the PR passed pre-merge verification criteria label Jul 27, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@bryan-cox: This PR has been marked as verified by @amasolov.

Details

In response to this:

/verified by @amasolov

See #8581 (comment)

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@bryan-cox

Copy link
Copy Markdown
Member

/ok-to-test

@openshift-ci

openshift-ci Bot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

@amasolov: all tests passed!

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@openshift-merge-bot
openshift-merge-bot Bot merged commit 02c0659 into openshift:main Jul 27, 2026
43 checks passed
@openshift-ci-robot

Copy link
Copy Markdown

@amasolov: Jira Issue Verification Checks: Jira Issue OCPBUGS-86494
✔️ This pull request was pre-merge verified.
✔️ All associated pull requests have merged.
✔️ All associated, merged pull requests were pre-merge verified.

Jira Issue OCPBUGS-86494 has been moved to the MODIFIED state and will move to the VERIFIED state when the change is available in an accepted nightly payload. 🕓

Details

In response to this:

What this PR does / why we need it:

On KubeVirt, the bootstrap userdata Secret in the HCP namespace (user-data-<nodepool>-<hash>) is shared by all VMs in a given NodePool generation. During a rolling update, HyperShift's secretJanitor and Token.cleanupOutdated() eagerly delete the old userdata Secret as soon as a new config version is computed. Any VMs from the previous generation that are still running will then fail CAPK reconciliation because their referenced Secret no longer exists.

This PR extends the existing AWS platform guard (which already preserves old userdata Secrets for a different reason) to also cover KubeVirt. With this change, old userdata Secrets are retained until the rollout finishes and old VMs are cleaned up.

Which issue(s) this PR fixes:

Fixes premature garbage collection of shared userdata Secrets on the KubeVirt provider during NodePool rolling updates.

Special notes for your reviewer:

  • The fix mirrors the existing AWS guard pattern. AWS preserves old userdata because of a CAPA bug (Do not return error if secret does not exist kubernetes-sigs/cluster-api-provider-aws#3805). KubeVirt needs it because the Secret is architecturally shared across VMs in a generation.
  • Long term, a more targeted cleanup (e.g. checking whether any old-generation VMs still reference the Secret) would be preferable, but this minimal change is safe and consistent with the existing pattern.

Manual testing

Verified on a real KubeVirt HostedCluster (OCP 4.20 management cluster with CNV 4.20.15, guest release 4.22.3, 2-replica NodePool). Triggered a rolling update by changing VM compute cores. The old userdata Secret was preserved throughout the rollout, both old machines were deleted cleanly, and the rollout completed with no errors. Full test evidence: #8581 (comment)

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Made with Cursor

Summary by CodeRabbit

  • Bug Fixes

  • Improved secret retention handling for KubeVirt platforms, ensuring user-data secrets are properly preserved during node pool operations and cluster updates.

  • Tests

  • Extended test coverage to validate secret lifecycle management behavior for KubeVirt platforms.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-merge-robot

Copy link
Copy Markdown
Contributor

Fix included in release 5.0.0-0.nightly-2026-07-28-081944

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. area/hypershift-operator Indicates the PR includes changes for the hypershift operator and API - outside an OCP release jira/severity-important Referenced Jira bug's severity is important for the branch this PR is targeting. jira/valid-bug Indicates that a referenced Jira bug is valid for the branch this PR is targeting. jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. lgtm Indicates that a PR is ready to be merged. ok-to-test Indicates a non-member PR verified by an org member that is safe to test. verified Signifies that the PR passed pre-merge verification criteria

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants