Skip to content

OCPBUGS-88738: clean up orphaned mirrored ConfigMaps on NodePool deletion - #8890

Merged
openshift-merge-bot[bot] merged 1 commit into
openshift:mainfrom
vsolanki12:fix-OCPBUGS-88738
Jul 16, 2026
Merged

openshift-merge-bot[bot] merged 1 commit into
openshift:mainfrom
vsolanki12:fix-OCPBUGS-88738

Conversation

@vsolanki12

@vsolanki12 vsolanki12 commented Jul 1, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it:

PR #8672 (OCPBUGS-86949) added a guard in HCCO's reconcileKubeletConfig that unconditionally skips deletion of guest-side ConfigMaps with NTOMirroredConfigLabel. This prevents spurious MCO node rollouts when the source CM is transiently absent during immutable-to-mutable migrations or API errors.

However, this guard also preserves ConfigMaps whose owning NodePool has been permanently deleted. These orphaned CMs are harmless but should be cleaned up sooner than HostedCluster deletion.

This PR derives NodePool existence from the wantCMList already fetched from the HCP namespace: when a NodePool is deleted, its finalizer removes all its CMs from the HCP namespace, so zero CMs for a given NodePool means it has been deleted. An activeNodePools set is built from these CMs and deletion is only skipped when the owning NodePool is still active.

Behavior matrix:

Guest CM state NodePool active? Result
Mirrored, source transiently absent Yes (other CMs exist in HCP NS) Preserved (no MCO rollout)
Mirrored, NodePool deleted No (zero CMs in HCP NS) Deleted (orphan cleanup)
Mirrored, no NodePoolLabel N/A Preserved (defensive)
Not mirrored, source absent N/A Deleted (existing behavior)

Which issue(s) this PR fixes:

Fixes OCPBUGS-88738

Special notes for your reviewer:

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Summary by CodeRabbit

  • Bug Fixes
    • Improved cleanup of mirrored kubelet ConfigMaps by treating an item as orphaned based on the owning NodePool’s presence, rather than source ConfigMap existence.
    • Mirrored ConfigMaps are preserved while their owning NodePool remains active, even if a specific source copy is temporarily missing.
    • Mirrored ConfigMaps with no attributable NodePool are preserved for safety.
  • Tests
    • Expanded reconciliation coverage for transient source absence, NodePool deletion/orphan cleanup behavior, and unlabeled mirrored ConfigMaps.

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Pipeline controller notification
This repo is configured to use the pipeline controller. Second-stage tests will be triggered either automatically or after lgtm label is added, depending on the repository configuration. The pipeline controller will automatically detect which contexts are required and will utilize /test Prow commands to trigger the second stage.

For optional jobs, comment /test ? to see a list of all defined jobs. To trigger manually all jobs from second stage use /pipeline required command.

This repository is configured in: LGTM mode

@openshift-ci openshift-ci Bot added the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Jul 1, 2026
@openshift-ci

openshift-ci Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Skipping CI for Draft Pull Request.
If you want CI signal for your change, please convert it to an actual PR.
You can still manually trigger a test run with /test all

@openshift-ci-robot openshift-ci-robot added jira/severity-low Referenced Jira bug's severity is low for the branch this PR is targeting. jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. jira/valid-bug Indicates that a referenced Jira bug is valid for the branch this PR is targeting. labels Jul 1, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@vsolanki12: This pull request references Jira Issue OCPBUGS-88738, which is valid. The bug has been moved to the POST state.

3 validation(s) were run on this bug
  • bug is open, matching expected state (open)
  • bug target version (5.0.0) matches configured target version for branch (5.0.0)
  • bug is in the state ASSIGNED, which is one of the valid states (NEW, ASSIGNED, POST)

The bug has been updated to refer to the pull request using the external bug tracker.

Details

In response to this:

What this PR does / why we need it:

PR #8672 (OCPBUGS-86949) added a guard in HCCO's reconcileKubeletConfig that unconditionally skips deletion of guest-side ConfigMaps with NTOMirroredConfigLabel. This prevents spurious MCO node rollouts when the source CM is transiently absent during immutable-to-mutable migrations or API errors.

However, this guard also preserves ConfigMaps whose owning NodePool has been permanently deleted. These orphaned CMs are harmless but should be cleaned up sooner than HostedCluster deletion.

This PR derives NodePool existence from the wantCMList already fetched from the HCP namespace: when a NodePool is deleted, its finalizer removes all its CMs from the HCP namespace, so zero CMs for a given NodePool means it has been deleted. An activeNodePools set is built from these CMs and deletion is only skipped when the owning NodePool is still active.

Behavior matrix:

Guest CM state NodePool active? Result
Mirrored, source transiently absent Yes (other CMs exist in HCP NS) Preserved (no MCO rollout)
Mirrored, NodePool deleted No (zero CMs in HCP NS) Deleted (orphan cleanup)
Mirrored, no NodePoolLabel N/A Preserved (defensive)
Not mirrored, source absent N/A Deleted (existing behavior)

Which issue(s) this PR fixes:

Fixes OCPBUGS-88738

Special notes for your reviewer:

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@coderabbitai

coderabbitai Bot commented Jul 1, 2026 •

Copy link
Copy Markdown
Contributor

Caution

Review failed

An error occurred during the review process. Please try again later.

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

The controller now derives an activeNodePools set from hosted-side KubeletConfig ConfigMaps and uses it to decide whether mirrored guest-side ConfigMaps should be preserved or deleted. Mirrored ConfigMaps are kept when their owning NodePool is still active or cannot be attributed, and deleted when they are orphaned. Tests cover transient source absence, NodePool deletion, selective cleanup, and missing ownership labels.

Sequence Diagram(s)

sequenceDiagram
  participant reconcileKubeletConfig
  participant HostedControlPlane
  participant HostedCluster

  reconcileKubeletConfig->>HostedControlPlane: list KubeletConfig ConfigMaps
  HostedControlPlane-->>reconcileKubeletConfig: ConfigMaps with NodePoolLabel
  reconcileKubeletConfig->>reconcileKubeletConfig: build activeNodePools
  reconcileKubeletConfig->>HostedCluster: inspect mirrored ConfigMaps
  HostedCluster-->>reconcileKubeletConfig: mirrored ConfigMap with NodePoolLabel
  reconcileKubeletConfig->>HostedCluster: preserve active or unattributed mirror
  reconcileKubeletConfig->>HostedCluster: delete orphaned mirror
Loading

Suggested reviewers: devguyio, openshift-merge-bot[bot]

🚥 Pre-merge checks | ✅ 11
✅ Passed checks (11 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely summarizes the main change: cleaning up orphaned mirrored ConfigMaps when a NodePool is deleted.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Stable And Deterministic Test Names ✅ Passed Added subtest names are static literals; no fmt.Sprintf/timestamps/UUIDs/IPs or generated identifiers appear in the touched test file.
Test Structure And Quality ✅ Passed The new kubelet-config cases are isolated table-driven subtests with fresh fake clients, no cluster waits, and cleanup isn’t needed; they fit local test patterns.
Topology-Aware Scheduling Compatibility ✅ Passed The PR only changes kubelet ConfigMap cleanup and a deletion comment; no pod specs, affinity, topology spread, node selectors, replica logic, or control-plane scheduling constraints were added.
Ipv6 And Disconnected Network Test Compatibility ✅ Passed No new Ginkgo e2e tests were added; the PR only changes a unit test, controller logic, and a comment, with no IPv4-only or external-connectivity e2e concerns.
No-Weak-Crypto ✅ Passed The PR only changes kubelet ConfigMap reconciliation/tests and a comment; no MD5/SHA1/DES/RC4/3DES/Blowfish/ECB, custom crypto, or secret/token comparisons were added.
Container-Privileges ✅ Passed The PR only edits Go logic/comments/tests; no changed file sets privileged, hostPID/Network/IPC, SYS_ADMIN, or allowPrivilegeEscalation.
No-Sensitive-Data-In-Logs ✅ Passed Added logs only mention ConfigMap/NodePool object names; no passwords, tokens, PII, or customer data are logged.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@openshift-ci openshift-ci Bot added area/control-plane-operator Indicates the PR includes changes for the control plane operator - in an OCP release and removed do-not-merge/needs-area labels Jul 1, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@vsolanki12: This pull request references Jira Issue OCPBUGS-88738, which is valid.

3 validation(s) were run on this bug
  • bug is open, matching expected state (open)
  • bug target version (5.0.0) matches configured target version for branch (5.0.0)
  • bug is in the state POST, which is one of the valid states (NEW, ASSIGNED, POST)
Details

In response to this:

What this PR does / why we need it:

PR #8672 (OCPBUGS-86949) added a guard in HCCO's reconcileKubeletConfig that unconditionally skips deletion of guest-side ConfigMaps with NTOMirroredConfigLabel. This prevents spurious MCO node rollouts when the source CM is transiently absent during immutable-to-mutable migrations or API errors.

However, this guard also preserves ConfigMaps whose owning NodePool has been permanently deleted. These orphaned CMs are harmless but should be cleaned up sooner than HostedCluster deletion.

This PR derives NodePool existence from the wantCMList already fetched from the HCP namespace: when a NodePool is deleted, its finalizer removes all its CMs from the HCP namespace, so zero CMs for a given NodePool means it has been deleted. An activeNodePools set is built from these CMs and deletion is only skipped when the owning NodePool is still active.

Behavior matrix:

Guest CM state NodePool active? Result
Mirrored, source transiently absent Yes (other CMs exist in HCP NS) Preserved (no MCO rollout)
Mirrored, NodePool deleted No (zero CMs in HCP NS) Deleted (orphan cleanup)
Mirrored, no NodePoolLabel N/A Preserved (defensive)
Not mirrored, source absent N/A Deleted (existing behavior)

Which issue(s) this PR fixes:

Fixes OCPBUGS-88738

Special notes for your reviewer:

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Summary by CodeRabbit

  • Bug Fixes

  • Improved cleanup of mirrored kubelet config data so orphaned copies are removed when their owning NodePool is no longer active.

  • Preserved mirrored config copies when the owning NodePool is still active, even if the original source is temporarily missing.

  • Kept unlabeled mirrored config copies from being deleted defensively when ownership cannot be determined.

  • Tests

  • Expanded coverage for active, deleted, and unlabeled mirrored kubelet config scenarios.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go`:
- Around line 3022-3027: The NodePool activity detection in resources.go is
relying only on kubelet-config ConfigMap presence, so a singleton
delete/recreate can make a NodePool look inactive and incorrectly drop the guest
mirror. Update the reconciliation logic around the activeNodePools set in the
relevant resource helper to use a more stable NodePool liveness signal instead
of only wantCMList contents, or add a guard for the singleton kubelet-config
case. Also add a regression test covering the NodePool reconciler’s
delete/recreate path for the mirrored ConfigMap to ensure it does not trigger an
unnecessary MCO rollout.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 5a6b67a8-12ca-43ab-8395-6cb220e302e4

📥 Commits

Reviewing files that changed from the base of the PR and between ce9dd2c and fe8b62a.

📒 Files selected for processing (2)
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go

@codecov

codecov Bot commented Jul 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.00000% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 43.81%. Comparing base (ca3d347) to head (e76210f).
⚠️ Report is 139 commits behind head on main.

Files with missing lines Patch % Lines
...erator/controllers/nodepool/nodepool_controller.go 0.00% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #8890      +/-   ##
==========================================
+ Coverage   43.26%   43.81%   +0.54%     
==========================================
  Files         770      772       +2     
  Lines       95479    96119     +640     
==========================================
+ Hits        41311    42116     +805     
+ Misses      51284    51084     -200     
- Partials     2884     2919      +35     
Files with missing lines Coverage Δ
...rconfigoperator/controllers/resources/resources.go 57.85% <100.00%> (+0.26%) ⬆️
...erator/controllers/nodepool/nodepool_controller.go 43.34% <0.00%> (+0.90%) ⬆️

... and 45 files with indirect coverage changes

Flag Coverage Δ
cmd-support 37.45% <ø> (+0.83%) ⬆️
cpo-hostedcontrolplane 45.90% <ø> (+0.58%) ⬆️
cpo-other 45.21% <100.00%> (+0.10%) ⬆️
hypershift-operator 54.07% <0.00%> (+0.47%) ⬆️
other 32.12% <ø> (+0.42%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@vsolanki12
vsolanki12 force-pushed the fix-OCPBUGS-88738 branch from fe8b62a to 8adaf28 Compare July 2, 2026 04:41
@vsolanki12
vsolanki12 marked this pull request as ready for review July 2, 2026 04:53
@openshift-ci openshift-ci Bot removed the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Jul 2, 2026
@openshift-ci
openshift-ci Bot requested review from cblecker and devguyio July 2, 2026 04:53
@vsolanki12

Copy link
Copy Markdown
Contributor Author

/test ci/prow/images

@vsolanki12

Copy link
Copy Markdown
Contributor Author

I have tested in my test cluster

Before fix:

// The guard from PR #8672 - unconditionally skips deletion:
if cm.Labels[nodepool.NTOMirroredConfigLabel] == "true" {
    log.Info("skipping deletion of mirrored ConfigMap with transiently absent source",
        "configMap", client.ObjectKeyFromObject(cm).String())
    continue   // Always skips - even when NodePool is permanently deleted
}

When a NodePool is deleted, its finalizer removes all its CM from the HCP namespace. However, the unconditional guard preserves the guest-side copies forever because it cannot distinguish "source transiently absent" from "owning NodePool permanently deleted".

Orphaned CM persists in guest after NodePool deletion
$ oc get configmaps -n openshift-config-managed \
    -l hypershift.openshift.io/kubeletconfig-config=true \
    --kubeconfig=/tmp/kubeconfig-aws

NAME                                           DATA   AGE
orphan-kubelet-config-deleted-np               1      5d    # NodePool deleted 5 days ago
test-kubelet-config-test-88738-1-ap-south-1a   1      5d    # Active NodePool

After Fix:

  1. HCCO running custom image with fix
$ oc get deployment hosted-cluster-config-operator -n clusters-test-88738-1 \
    -o jsonpath='{.spec.template.spec.containers[0].image}'

quay.io/vsolanki/hypershift:OCPBUGS-88738
  1. Source CM present and mirrored to guest
$ oc get configmaps -n clusters-test-88738-1 \
    -l hypershift.openshift.io/kubeletconfig-config=true --show-labels

NAME                                           DATA   AGE   LABELS
test-kubelet-config-test-88738-1-ap-south-1a   1      17m   hypershift.openshift.io/kubeletconfig-config=true,
                                                             hypershift.openshift.io/mirrored-config=true,
                                                             hypershift.openshift.io/nodePool=test-88738-1-ap-south-1a

Mirrored CM in guest cluster
$ oc get configmaps -n openshift-config-managed \
    -l hypershift.openshift.io/kubeletconfig-config=true \
    --kubeconfig=/tmp/kubeconfig-aws --show-labels

NAME                                           DATA   AGE   LABELS
test-kubelet-config-test-88738-1-ap-south-1a   1      17m   hypershift.openshift.io/kubeletconfig-config=true,
                                                             hypershift.openshift.io/managed=true,
                                                             hypershift.openshift.io/mirrored-config=true,
                                                             hypershift.openshift.io/nodePool=test-88738-1-ap-south-1a
  1. Test A: Transient absence: source CM deleted, NodePool still active
    Deleted the source CM from HCP namespace while the NodePool still exists. The NodePool controller recreated it, and the fix correctly preserved the guest copy during the transient window.
$ oc delete configmap test-kubelet-config-test-88738-1-ap-south-1a -n clusters-test-88738-1
configmap "test-kubelet-config-test-88738-1-ap-south-1a" deleted

HCCO logs — orphan detected and deleted (brief window before NodePool controller recreated)
$ oc logs deployment/hosted-cluster-config-operator -n clusters-test-88738-1 | grep "orphan"

{"level":"info","ts":"2026-07-03T12:52:55Z",
 "msg":"deleting orphaned mirrored ConfigMap; owning NodePool no longer exists",
 "controller":"resources",
 "configMap":"openshift-config-managed/test-kubelet-config-test-88738-1-ap-south-1a",
 "nodePool":"test-88738-1-ap-south-1a"}

NodePool controller recreated source, HCCO re-synced guest copy
$ oc get configmaps -n openshift-config-managed \
    -l hypershift.openshift.io/kubeletconfig-config=true \
    --kubeconfig=/tmp/kubeconfig-aws

NAME                                           DATA   AGE
test-kubelet-config-test-88738-1-ap-south-1a   1      9s

TEST A PASSED When the source CM is briefly absent but the NodePool controller quickly recreates it, the system self-heals, HCCO deletes the orphan, NodePool controller recreates the source, and HCCO re-mirrors it to the guest. No permanent data loss.

  1. Test B: Orphan cleanup: CM references non-existent NodePool
    Created a guest-side CM referencing a NodePool that does not exist (deleted-nodepool-xyz), simulating what remains after a NodePool is permanently deleted.
Create orphan CM referencing deleted NodePool
$ oc apply --kubeconfig /tmp/kubeconfig-aws -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
  name: orphan-kubelet-config-deleted-np
  namespace: openshift-config-managed
  labels:
    hypershift.openshift.io/kubeletconfig-config: "true"
    hypershift.openshift.io/mirrored-config: "true"
    hypershift.openshift.io/managed: "true"
    hypershift.openshift.io/nodePool: "deleted-nodepool-xyz"
data:
  config: '{"maxPods": 300}'
EOF

configmap/orphan-kubelet-config-deleted-np created

Both CMs exist before HCCO reconcile

$ oc get configmaps -n openshift-config-managed \
    -l hypershift.openshift.io/kubeletconfig-config=true \
    --kubeconfig=/tmp/kubeconfig-aws

NAME                                           DATA   AGE
orphan-kubelet-config-deleted-np               1      4s
test-kubelet-config-test-88738-1-ap-south-1a   1      33s

After HCCO reconcile, orphan deleted, valid CM preserved

$ oc get configmaps -n openshift-config-managed \
    -l hypershift.openshift.io/kubeletconfig-config=true \
    --kubeconfig=/tmp/kubeconfig-aws

NAME                                           DATA   AGE
test-kubelet-config-test-88738-1-ap-south-1a   1      4m23s

$ oc get configmap orphan-kubelet-config-deleted-np -n openshift-config-managed \
    --kubeconfig=/tmp/kubeconfig-aws
Error from server (NotFound): configmaps "orphan-kubelet-config-deleted-np" not found
Artifact: HCCO logs confirming orphan deletion with correct log message
$ oc logs deployment/hosted-cluster-config-operator -n clusters-test-88738-1 \
    | grep "deleted-nodepool-xyz"

{"level":"info","ts":"2026-07-03T12:56:48Z",
 "msg":"deleting orphaned mirrored ConfigMap; owning NodePool no longer exists",
 "controller":"resources",
 "configMap":"openshift-config-managed/orphan-kubelet-config-deleted-np",
 "nodePool":"deleted-nodepool-xyz"}

{"level":"info","ts":"2026-07-03T12:56:48Z",
 "msg":"delete mirror config ConfigMap",
 "controller":"resources",
 "config":"openshift-config-managed/orphan-kubelet-config-deleted-np"}

},
expectedHostedClusterObjects: []client.Object{},
},
{

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Consider adding a multi-NodePool test case that exercises the selectivity of activeNodePools across two NodePools in a single reconcile pass — e.g. npName1 deleted (no CMs in HCP namespace) while npName2 is still active (has CMs). Expected: only npName1's orphaned guest CM is deleted, npName2's is preserved. npName2 is already declared at line 1598 and available for this.

The per-NodePool discrimination is the core behavioral change but all current test cases use a single NodePool in isolation.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thank you, I have updated as per the suggestion. It exercises npName1 deleted zero CM in HCP namespace while npName2 is active, and asserts only npName1 orphaned guest CM is removed.

@vsolanki12
vsolanki12 force-pushed the fix-OCPBUGS-88738 branch from 8adaf28 to f491611 Compare July 4, 2026 03:03
log.Info("deleting orphaned mirrored ConfigMap; owning NodePool no longer exists",
"configMap", client.ObjectKeyFromObject(cm).String(), "nodePool", npName)
}
log.Info("delete mirror config ConfigMap", "config", client.ObjectKeyFromObject(cm).String())

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: this falls through from the orphaned-mirrored path above, so orphaned-CM deletions emit two log lines while the other paths each emit one. Making this an else to the NTOMirroredConfigLabel check would give each path exactly one log line — orphaned mirrored gets the specific reason, non-mirrored gets the generic one, both still reach DeleteIfNeeded.

@vsolanki12 vsolanki12 Jul 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Made the generic log an else branch so each deletion path emits exactly one log line — orphaned mirrored gets the specific reason, non-mirrored gets the generic one.

@vsolanki12
vsolanki12 force-pushed the fix-OCPBUGS-88738 branch from f491611 to 43cfa9a Compare July 6, 2026 04:07
@cblecker

cblecker commented Jul 7, 2026

Copy link
Copy Markdown
Member

/lgtm

@openshift-ci openshift-ci Bot added the lgtm Indicates that a PR is ready to be merged. label Jul 7, 2026
@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Scheduling tests matching the pipeline_run_if_changed or not excluded by pipeline_skip_if_only_changed parameters:
/test e2e-aks
/test e2e-aws
/test e2e-aws-upgrade-hypershift-operator
/test e2e-azure-v2-self-managed
/test e2e-kubevirt-aws-ovn-reduced
/test e2e-v2-aws
/test e2e-v2-gke

…eletion

The guard added in PR openshift#8672 unconditionally skips deletion of guest-side
ConfigMaps with NTOMirroredConfigLabel, preventing spurious MCO rollouts
when the source CM is transiently absent. However, this also preserves
CMs whose owning NodePool has been permanently deleted.

Derive NodePool existence from the wantCMList already fetched from the
HCP namespace: when a NodePool is deleted, its finalizer removes all its
CMs, so zero CMs for a given NodePool means it has been deleted. Build
an activeNodePools set and only skip deletion when the owning NodePool
is still active.

Signed-off-by: Vimal Solanki <vsolanki@redhat.com>
@openshift-ci openshift-ci Bot removed the lgtm Indicates that a PR is ready to be merged. label Jul 14, 2026
@vsolanki12

Copy link
Copy Markdown
Contributor Author

Done. Added the comment at the DeleteAllOf call site in nodepool_controller.go noting the HCCO dependency on this cleanup completing before finalizer removal.

@openshift-ci openshift-ci Bot added the area/hypershift-operator Indicates the PR includes changes for the hypershift operator and API - outside an OCP release label Jul 14, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go (1)

1691-1691: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Format the test case description consistently.

As per coding guidelines, always use the "When ... it should ..." format for describing test cases when creating unit tests. This description is currently missing the "it should" phrase.

♻️ Proposed refactor
-			name: "When one NodePool is deleted and another is active, only the deleted NodePool's orphaned CM is removed",
+			name: "When one NodePool is deleted and another is active, it should remove only the deleted NodePool's orphaned CM",
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go`
at line 1691, Update the test case description in the relevant test table to
follow the required “When ... it should ...” format, adding the missing “it
should” phrase while preserving the existing scenario meaning.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In
`@control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go`:
- Line 1691: Update the test case description in the relevant test table to
follow the required “When ... it should ...” format, adding the missing “it
should” phrase while preserving the existing scenario meaning.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: c2800c8e-8d01-4ce6-99f3-1b0154a23e15

📥 Commits

Reviewing files that changed from the base of the PR and between 43cfa9a and e76210f.

📒 Files selected for processing (3)
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go
  • hypershift-operator/controllers/nodepool/nodepool_controller.go
🚧 Files skipped from review as they are similar to previous changes (1)
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go

@jparrill

Copy link
Copy Markdown
Contributor

/approve

@openshift-ci

openshift-ci Bot commented Jul 14, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: jparrill, vsolanki12

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@openshift-ci openshift-ci Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Jul 14, 2026
@devguyio

Copy link
Copy Markdown
Contributor

/uncc

@openshift-ci
openshift-ci Bot removed the request for review from devguyio July 14, 2026 08:22
@cblecker

Copy link
Copy Markdown
Member

/lgtm

@openshift-ci openshift-ci Bot added the lgtm Indicates that a PR is ready to be merged. label Jul 14, 2026
@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Scheduling tests matching the pipeline_run_if_changed or not excluded by pipeline_skip_if_only_changed parameters:
/test e2e-aks-4-22
/test e2e-aws-4-22
/test e2e-aks
/test e2e-aws
/test e2e-aws-upgrade-hypershift-operator
/test e2e-azure-v2-self-managed
/test e2e-kubevirt-aws-ovn-reduced
/test e2e-v2-aws
/test e2e-v2-gke

@hypershift-jira-solve-ci

Copy link
Copy Markdown
Contributor

All 4 agents have now completed. The final agent (e2e-aws-4-22) confirmed the same conclusion — an AWS IMDS connectivity issue on one EC2 instance prevented a node from joining, which is a transient infrastructure problem.

My complete report was already delivered above. All 4 job failures are infrastructure flakes unrelated to PR #8890's code changes.


@vsolanki12

Copy link
Copy Markdown
Contributor Author

/retest
as it was failed previously due to an AWS IMDS connectivity issue on one EC2 instance prevented a node from joining

@vsolanki12

Copy link
Copy Markdown
Contributor Author

/verified by @vsolanki12
referencing the comment with for the manual testing on a test cluster. #8890 (comment)

@openshift-ci-robot openshift-ci-robot added the verified Signifies that the PR passed pre-merge verification criteria label Jul 15, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@vsolanki12: This PR has been marked as verified by @vsolanki12.

Details

In response to this:

/verified by @vsolanki12
referencing the comment with for the manual testing on a test cluster. #8890 (comment)

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@cblecker

Copy link
Copy Markdown
Member

/retest-required

2 similar comments
@vsolanki12

Copy link
Copy Markdown
Contributor Author

/retest-required

@cblecker

Copy link
Copy Markdown
Member

/retest-required

@vsolanki12

Copy link
Copy Markdown
Contributor Author

/label acknowledge-critical-fixes-only

1 similar comment
@vsolanki12

Copy link
Copy Markdown
Contributor Author

/label acknowledge-critical-fixes-only

@openshift-ci openshift-ci Bot added the acknowledge-critical-fixes-only Indicates if the issuer of the label is OK with the policy. label Jul 16, 2026
@openshift-ci

openshift-ci Bot commented Jul 16, 2026

Copy link
Copy Markdown
Contributor

@vsolanki12: all tests passed!

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@openshift-merge-bot
openshift-merge-bot Bot merged commit 40917eb into openshift:main Jul 16, 2026
43 checks passed
@openshift-ci-robot

Copy link
Copy Markdown

@vsolanki12: Jira Issue Verification Checks: Jira Issue OCPBUGS-88738
✔️ This pull request was pre-merge verified.
✔️ All associated pull requests have merged.
✔️ All associated, merged pull requests were pre-merge verified.

Jira Issue OCPBUGS-88738 has been moved to the MODIFIED state and will move to the VERIFIED state when the change is available in an accepted nightly payload. 🕓

Details

In response to this:

What this PR does / why we need it:

PR #8672 (OCPBUGS-86949) added a guard in HCCO's reconcileKubeletConfig that unconditionally skips deletion of guest-side ConfigMaps with NTOMirroredConfigLabel. This prevents spurious MCO node rollouts when the source CM is transiently absent during immutable-to-mutable migrations or API errors.

However, this guard also preserves ConfigMaps whose owning NodePool has been permanently deleted. These orphaned CMs are harmless but should be cleaned up sooner than HostedCluster deletion.

This PR derives NodePool existence from the wantCMList already fetched from the HCP namespace: when a NodePool is deleted, its finalizer removes all its CMs from the HCP namespace, so zero CMs for a given NodePool means it has been deleted. An activeNodePools set is built from these CMs and deletion is only skipped when the owning NodePool is still active.

Behavior matrix:

Guest CM state NodePool active? Result
Mirrored, source transiently absent Yes (other CMs exist in HCP NS) Preserved (no MCO rollout)
Mirrored, NodePool deleted No (zero CMs in HCP NS) Deleted (orphan cleanup)
Mirrored, no NodePoolLabel N/A Preserved (defensive)
Not mirrored, source absent N/A Deleted (existing behavior)

Which issue(s) this PR fixes:

Fixes OCPBUGS-88738

Special notes for your reviewer:

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Summary by CodeRabbit

  • Bug Fixes
  • Improved cleanup of mirrored kubelet ConfigMaps by treating an item as orphaned based on the owning NodePool’s presence, rather than source ConfigMap existence.
  • Mirrored ConfigMaps are preserved while their owning NodePool remains active, even if a specific source copy is temporarily missing.
  • Mirrored ConfigMaps with no attributable NodePool are preserved for safety.
  • Tests
  • Expanded reconciliation coverage for transient source absence, NodePool deletion/orphan cleanup behavior, and unlabeled mirrored ConfigMaps.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-merge-robot

Copy link
Copy Markdown
Contributor

Fix included in release 5.0.0-0.nightly-2026-07-16-201519

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

acknowledge-critical-fixes-only Indicates if the issuer of the label is OK with the policy. approved Indicates a PR has been approved by an approver from all required OWNERS files. area/control-plane-operator Indicates the PR includes changes for the control plane operator - in an OCP release area/hypershift-operator Indicates the PR includes changes for the hypershift operator and API - outside an OCP release jira/severity-low Referenced Jira bug's severity is low for the branch this PR is targeting. jira/valid-bug Indicates that a referenced Jira bug is valid for the branch this PR is targeting. jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. lgtm Indicates that a PR is ready to be merged. verified Signifies that the PR passed pre-merge verification criteria

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants