Skip to content

OCPBUGS-84368: add safe-to-evict annotation and remove tolerations to fix autoscaler scale-down and drain loop - #8338

Merged
openshift-merge-bot[bot] merged 6 commits into
openshift:mainfrom
sdminonne:fix/kas-connection-checker-pdb-safe-to-evict
Sep 5, 2026
Merged

openshift-merge-bot[bot] merged 6 commits into
openshift:mainfrom
sdminonne:fix/kas-connection-checker-pdb-safe-to-evict

Conversation

@sdminonne

@sdminonne sdminonne commented Apr 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Problem

The kas-connection-checker deployment runs in the kube-system namespace without a PDB or safe-to-evict annotation. The cluster autoscaler treats pods in kube-system without a PDB as system-critical pods that cannot be evicted during scale-down, blocking node draining on underutilized nodes. See also the autoscaler system drainability rule.

Additionally, the deployment used a blanket {Effect: NoSchedule, Operator: Exists} toleration that matched the cordon taint (node.kubernetes.io/unschedulable:NoSchedule), causing replacement pods to be scheduled back onto cordoned nodes during drain — creating a destructive evict-reschedule loop.

Solution

  1. Add cluster-autoscaler.kubernetes.io/safe-to-evict: "true" annotation to the pod template. This bypasses the autoscaler's system drainability rule for kube-system pods, allowing nodes to be scaled down. The annotation is preferred over a PDB because it doesn't block scale-to-zero operations.

  2. Remove all custom tolerations. The previous blanket NoSchedule toleration matched the cordon taint, causing the evict-reschedule loop described above. In a HyperShift hosted cluster there are no master nodes in the guest cluster, and the checker doesn't need to run on infra nodes, so no NoSchedule tolerations are needed. Kubernetes provides default NoExecute tolerations (300s grace period) for unreachable and not-ready via the DefaultTolerationSeconds admission controller.

  3. Add TopologySpreadConstraint with ScheduleAnyway to spread the 3 replicas across nodes. Using ScheduleAnyway (soft constraint) avoids blocking scheduling on small clusters (1-2 nodes).

  4. Preserve existing annotations and labels during reconciliation instead of replacing the entire map. Explicitly delete the stale openshift.io/required-scc annotation from older HCCO versions (kube-system is exempt from SCC admission via the systemNamespaces allowlist in origin).

Scope and impact

The tolerations are only removed from the kas-connection-checker Deployment in the guest cluster's kube-system namespace — no other workloads are affected. If all worker nodes carry custom NoSchedule taints and the checker pods cannot schedule, the ControlPlaneConnectionAvailable condition will report False (not Unknown) with reason ControlPlaneConnectionKASAccessFailedReason, because the ConfigMap never receives a lastSucceeded timestamp. This is an accurate signal: if the checker cannot run, data-plane-to-control-plane connectivity genuinely cannot be verified. In the normal case, the 3 replicas with ScheduleAnyway topology spread will schedule on any untainted nodes.

Changes

  • control-plane-operator/.../resources/resources.go: Add safe-to-evict annotation to pod template (merge, don't replace); remove stale required-scc annotation; remove all custom tolerations; add TopologySpreadConstraint; broaden log message from "deployment" to "resources".
  • control-plane-operator/.../resources/resources_test.go: Update tests to validate the safe-to-evict annotation, assert no custom tolerations, verify topology spread constraint. Rename test to TestReconcileKASConnectionChecker per TESTING.md convention.
  • test/e2e/v2/tests/hosted_cluster_compliance_test.go: Add EnsureKASConnectionCheckerSpecTest to the v2 e2e suite, gated behind VersionAtLeast(Version423).
  • test/e2e/create_cluster_test.go: Remove EnsureKASConnectionCheckerSpec call (moved to v2).
  • test/e2e/util/util.go: Remove EnsureKASConnectionCheckerSpec helper (moved to v2).

Test plan

  • go test ./control-plane-operator/hostedclusterconfigoperator/controllers/resources/... -count=1 passes
  • make verify passes
  • Verify on a live cluster that the annotation is present on kas-connection-checker pods
  • Verify cluster autoscaler can scale down nodes running kas-connection-checker pods

Fixes: https://redhat.atlassian.net/browse/OCPBUGS-84368

🤖 Generated with Claude Code

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Pipeline controller notification
This repo is configured to use the pipeline controller. Second-stage tests will be triggered either automatically or after lgtm label is added, depending on the repository configuration. The pipeline controller will automatically detect which contexts are required and will utilize /test Prow commands to trigger the second stage.

For optional jobs, comment /test ? to see a list of all defined jobs. To trigger manually all jobs from second stage use /pipeline required command.

This repository is configured in: LGTM mode

@coderabbitai

coderabbitai Bot commented Apr 26, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

This change adds PodDisruptionBudget (PDB) support to the KAS connection checker component. A new helper function constructs PDB manifests with appropriate metadata. The reconciliation flow for the KAS connection checker is updated to annotate the Deployment pod template as safe to evict and then create or update a PodDisruptionBudget with MinAvailable=1, a selector matching KAS connection checker pods, and unhealthy pod eviction enabled. A new constant defines the cluster autoscaler safe-to-evict annotation key. Test coverage is expanded to validate PDB reconciliation behavior, including edge cases and error scenarios.

Sequence Diagram

sequenceDiagram
    participant Reconciler as Reconciler<br/>(KAS Connection Checker)
    participant K8sAPI as Kubernetes API
    participant Deployment as Deployment<br/>Resource
    participant PDB as PodDisruptionBudget<br/>Resource
    
    Reconciler->>K8sAPI: Fetch existing Deployment
    K8sAPI-->>Reconciler: Return Deployment
    
    Reconciler->>Deployment: Annotate pod template<br/>(safe-to-evict)
    Reconciler->>K8sAPI: Create/Update Deployment
    K8sAPI-->>Reconciler: Deployment reconciled
    
    Reconciler->>K8sAPI: Fetch existing PDB
    K8sAPI-->>Reconciler: Return PDB (or none)
    
    Reconciler->>PDB: Set minAvailable=1<br/>Set selector (app=KASConnectionChecker)<br/>Set unhealthyPodEvictionPolicy=AlwaysAllow
    Reconciler->>K8sAPI: Create/Update PDB
    K8sAPI-->>Reconciler: PDB reconciled
    
    Reconciler-->>Reconciler: Return reconciliation result
Loading
🚥 Pre-merge checks | ✅ 9 | ❌ 3

❌ Failed checks (2 warnings, 1 inconclusive)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 20.83% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
Topology-Aware Scheduling Compatibility ⚠️ Warning Fixed 3-replica deployment lacks topology awareness for SNO, Two-Node, and HyperShift topologies, causing scheduling failures. Implement topology-aware logic checking infrastructure.Status.ControlPlaneTopology to adjust replica count and PDB settings per cluster topology.
Test Structure And Quality ❓ Inconclusive PR summary indicates resources_test.go uses Go testing package (t.Parallel() removal), not Ginkgo, making the Ginkgo-specific check inapplicable without direct file inspection. Verify the test framework used in resources_test.go to determine if Ginkgo assessment criteria apply or if alternative testing package standards should be used.
✅ Passed checks (9 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Stable And Deterministic Test Names ✅ Passed All test names in modified files are stable, static string constants with no dynamic elements like timestamps, UUIDs, pod names, node names, or namespace suffixes embedded in test names.
Microshift Test Compatibility ✅ Passed This PR does not add any new Ginkgo e2e tests; it only modifies standard Go unit tests in resources_test.go and includes a minor formatting change in nodepool_test.go.
Single Node Openshift (Sno) Test Compatibility ✅ Passed This PR does not add new Ginkgo e2e tests. Changes are unit tests using standard Go testing package, not Ginkgo framework.
Ote Binary Stdout Contract ✅ Passed The pull request does not violate the OTE Binary Stdout Contract. All modifications use only Kubernetes API calls and controller-runtime logging.
Ipv6 And Disconnected Network Test Compatibility ✅ Passed The PR does not add any new Ginkgo e2e tests using patterns like It(), Describe(), Context(), or When(). New test code is in a unit test file using standard Go testing.T framework.
Title check ✅ Passed The title directly addresses the main objective of the PR: adding a safe-to-evict annotation and addressing autoscaler scale-down and drain loop issues for the KAS connection checker.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@openshift-ci openshift-ci Bot added area/control-plane-operator Indicates the PR includes changes for the control plane operator - in an OCP release and removed do-not-merge/needs-area labels Apr 26, 2026
@sdminonne
sdminonne force-pushed the fix/kas-connection-checker-pdb-safe-to-evict branch from 67a200b to 238cc6a Compare April 26, 2026 10:56
@openshift-ci openshift-ci Bot added the area/testing Indicates the PR includes changes for e2e testing label Apr 26, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go (1)

2841-2855: Add an existing-PDB reconciliation case.

These assertions only cover the create path. Since the controller now reconciles the PDB on every run, please add a case that seeds a kas-connection-checker PDB with maxUnavailable and verifies reconcile rewrites it to minAvailable=1/AlwaysAllow.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go`
around lines 2841 - 2855, Add a new test case that seeds an existing
PodDisruptionBudget named by manifests.KASConnectionCheckerName in namespace
manifests.KASConnectionCheckerNamespace with Spec.MaxUnavailable set (e.g., 1)
and UnhealthyPodEvictionPolicy set to policyv1.DoNotEvict, then invoke the
controller reconcile path used by other tests (the same reconcile helper used in
resources_test.go) and assert the PDB is mutated: fetch the PDB via c.Get and
verify pdb.Spec.MinAvailable equals intstr.FromInt32(1),
pdb.Spec.UnhealthyPodEvictionPolicy equals policyv1.AlwaysAllow, and
pdb.Spec.Selector still matches app=manifests.KASConnectionCheckerName to
confirm reconciliation rewrote maxUnavailable to minAvailable and enforced
AlwaysAllow.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In
`@control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go`:
- Around line 1756-1764: The update closure passed to r.CreateOrUpdate for the
PodDisruptionBudget (pdb) should clear any existing pdb.Spec.MaxUnavailable
before assigning MinAvailable, because PodDisruptionBudgetSpec rejects objects
with both fields set; inside the anonymous func (the CreateOrUpdate callback
that mutates pdb), explicitly set pdb.Spec.MaxUnavailable = nil (or reset it)
prior to setting pdb.Spec.MinAvailable = ptr.To(intstr.FromInt32(1)) and
pdb.Spec.UnhealthyPodEvictionPolicy, referencing the pdb variable and the
CreateOrUpdate call that manages the kas-connection-checker PDB
(manifests.KASConnectionCheckerName).

---

Nitpick comments:
In
`@control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go`:
- Around line 2841-2855: Add a new test case that seeds an existing
PodDisruptionBudget named by manifests.KASConnectionCheckerName in namespace
manifests.KASConnectionCheckerNamespace with Spec.MaxUnavailable set (e.g., 1)
and UnhealthyPodEvictionPolicy set to policyv1.DoNotEvict, then invoke the
controller reconcile path used by other tests (the same reconcile helper used in
resources_test.go) and assert the PDB is mutated: fetch the PDB via c.Get and
verify pdb.Spec.MinAvailable equals intstr.FromInt32(1),
pdb.Spec.UnhealthyPodEvictionPolicy equals policyv1.AlwaysAllow, and
pdb.Spec.Selector still matches app=manifests.KASConnectionCheckerName to
confirm reconciliation rewrote maxUnavailable to minAvailable and enforced
AlwaysAllow.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: dff10846-33f5-40b0-99c8-a423c96284f5

📥 Commits

Reviewing files that changed from the base of the PR and between 67a200b and 238cc6a.

📒 Files selected for processing (4)
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/manifests/kasconnectionchecker.go
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go
  • test/e2e/nodepool_test.go
✅ Files skipped from review due to trivial changes (1)
  • test/e2e/nodepool_test.go

@openshift-ci openshift-ci Bot added the area/hypershift-operator Indicates the PR includes changes for the hypershift operator and API - outside an OCP release label Apr 26, 2026
@sdminonne
sdminonne force-pushed the fix/kas-connection-checker-pdb-safe-to-evict branch from 784e934 to cbad894 Compare April 26, 2026 12:02
@codecov

codecov Bot commented Apr 26, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 47.01%. Comparing base (8bc8333) to head (a556ce2).
⚠️ Report is 31 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #8338      +/-   ##
==========================================
+ Coverage   46.99%   47.01%   +0.01%     
==========================================
  Files         786      786              
  Lines       99118    99132      +14     
==========================================
+ Hits        46579    46602      +23     
+ Misses      49390    49382       -8     
+ Partials     3149     3148       -1     
Files with missing lines Coverage Δ
...rconfigoperator/controllers/resources/resources.go 57.88% <100.00%> (+0.02%) ⬆️

... and 1 file with indirect coverage changes

Flag Coverage Δ
cmd-support 40.50% <ø> (+0.05%) ⬆️
cpo-hostedcontrolplane 50.32% <ø> (ø)
cpo-other 47.61% <100.00%> (+<0.01%) ⬆️
hypershift-operator 57.24% <ø> (ø)
other 34.70% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@sdminonne
sdminonne force-pushed the fix/kas-connection-checker-pdb-safe-to-evict branch from cbad894 to 8979ea8 Compare April 26, 2026 15:53

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go (1)

580-582: Clarify the top-level wrapped error message.

This path now reconciles both Deployment and PDB, but the wrapper still says “deployment” only. A broader message will make failure triage clearer.

Proposed tweak
-            errs = append(errs, fmt.Errorf("failed to reconcile KAS connection checker deployment: %w", err))
+            errs = append(errs, fmt.Errorf("failed to reconcile KAS connection checker resources: %w", err))
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go`
around lines 580 - 582, The current error wrapper in the call that invokes
reconcileKASConnectionChecker uses "failed to reconcile KAS connection checker
deployment" but that function reconciles both the Deployment and the
PodDisruptionBudget; update the wrapped message to reflect both resources (e.g.,
"failed to reconcile KAS connection checker resources (deployment and PDB)" or
similar) so failures from reconcileKASConnectionChecker are not
misleading—change the fmt.Errorf wrapper string in the block that calls
reconcileKASConnectionChecker(ctx, hcp, cliImage) accordingly.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In
`@control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go`:
- Around line 580-582: The current error wrapper in the call that invokes
reconcileKASConnectionChecker uses "failed to reconcile KAS connection checker
deployment" but that function reconciles both the Deployment and the
PodDisruptionBudget; update the wrapped message to reflect both resources (e.g.,
"failed to reconcile KAS connection checker resources (deployment and PDB)" or
similar) so failures from reconcileKASConnectionChecker are not
misleading—change the fmt.Errorf wrapper string in the block that calls
reconcileKASConnectionChecker(ctx, hcp, cliImage) accordingly.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 4e5098fd-38b6-4599-929b-7eba3028cb17

📥 Commits

Reviewing files that changed from the base of the PR and between cbad894 and 8979ea8.

📒 Files selected for processing (5)
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/manifests/kasconnectionchecker.go
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go
  • support/config/types.go
  • test/e2e/nodepool_test.go
✅ Files skipped from review due to trivial changes (1)
  • test/e2e/nodepool_test.go
🚧 Files skipped from review as they are similar to previous changes (1)
  • control-plane-operator/hostedclusterconfigoperator/controllers/resources/manifests/kasconnectionchecker.go

@sdminonne

Copy link
Copy Markdown
Contributor Author

/test e2e-conformance

@sdminonne sdminonne changed the title fix(kas-connection-checker): add PDB and safe-to-evict annotation OCPBUGS-84368: add PDB and safe-to-evict annotation to kas-connection-checker Apr 27, 2026
@openshift-ci-robot openshift-ci-robot added jira/severity-important Referenced Jira bug's severity is important for the branch this PR is targeting. jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. jira/valid-bug Indicates that a referenced Jira bug is valid for the branch this PR is targeting. labels Apr 27, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@sdminonne: This pull request references Jira Issue OCPBUGS-84368, which is valid. The bug has been moved to the POST state.

3 validation(s) were run on this bug
  • bug is open, matching expected state (open)
  • bug target version (5.0.0) matches configured target version for branch (5.0.0)
  • bug is in the state New, which is one of the valid states (NEW, ASSIGNED, POST)

The bug has been updated to refer to the pull request using the external bug tracker.

Details

In response to this:

Summary

Problem

The kas-connection-checker deployment in kube-system uses system-node-critical priority class, which prevents the cluster autoscaler from evicting its pods during scale-down. This blocks node draining and prevents efficient cluster scale-down when kas-connection-checker pods are running on underutilized nodes.

Solution

  • Add cluster-autoscaler.kubernetes.io/safe-to-evict: "true" annotation to the kas-connection-checker pod template so the cluster autoscaler can evict pods despite system-node-critical priority class.
  • Add a PodDisruptionBudget with minAvailable: 1 and unhealthyPodEvictionPolicy: AlwaysAllow to guarantee at least one replica remains available during voluntary disruptions while still allowing scale-down.
  • Extract the cluster-autoscaler.kubernetes.io/safe-to-evict annotation key to a shared constant (PodSafeToEvictKey) in support/config alongside the existing PodSafeToEvictLocalVolumesKey.
  • Rename reconcileKASConnectionCheckerDeployment to reconcileKASConnectionChecker since it now manages the PDB in addition to the deployment.

Changes

  • support/config/types.go: New PodSafeToEvictKey constant.
  • control-plane-operator/.../manifests/kasconnectionchecker.go: New KASConnectionCheckerPodDisruptionBudget() manifest builder.
  • control-plane-operator/.../resources/resources.go: Add safe-to-evict annotation to pod template; create/reconcile PDB with minAvailable: 1 and AlwaysAllow eviction policy; clear any stale maxUnavailable field.
  • control-plane-operator/.../resources/resources_test.go: New test cases for PDB creation, PDB reconciliation from maxUnavailable to minAvailable, PDB creation failure error propagation, and safe-to-evict annotation assertions on both create and update paths.
  • test/e2e/nodepool_test.go: Minor string concatenation formatting fix.

Test plan

  • go test ./control-plane-operator/hostedclusterconfigoperator/controllers/resources/ -run Test_reconciler_reconcileKASConnectionChecker -v passes
  • go test ./control-plane-operator/hostedclusterconfigoperator/controllers/resources/... passes (including TestReconcileErrorHandling and PDB failure injection)
  • Verify on a live cluster that the PDB is created in kube-system with correct spec
  • Verify cluster autoscaler can scale down nodes running kas-connection-checker pods

🤖 Generated with Claude Code

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@csrwng

csrwng commented Apr 27, 2026

Copy link
Copy Markdown
Contributor

/approve

@openshift-ci

openshift-ci Bot commented Apr 27, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: csrwng, sdminonne

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@openshift-ci openshift-ci Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Apr 27, 2026
@sdminonne

Copy link
Copy Markdown
Contributor Author

According to claude

  Summary: This is a KNOWN FLAKY upstream Kubernetes e2e test, unrelated to the PR #8338 changes.
           The test creates a CRD with a custom ResourceQuota and waits for the quota status to
           reflect the custom resource lifecycle. It timed out after ~61 seconds with
           "context deadline exceeded" at resource_quota.go:683. PR #8338 only modifies
           kas-connection-checker (PDB + safe-to-evict annotation) and has no interaction
           with ResourceQuota or CRD lifecycle machinery. The test was not eligible for retries
           per the retry strategy. Overall suite: 1929 passed, 1 blocking fail, 1 informing fail,
           1 flaky, 2127 skipped.

  Evidence:
    - JUnit XML failure message:
      <failure>fail [k8s.io/kubernetes/test/e2e/apimachinery/resource_quota.go:683]: context deadline exceeded</failure>
    - Test output from JUnit system-out:
      "Creating a Custom Resource Definition" at 20:29:39
      "Truncated CRD plural name from 'e2e-test-e2e-resourcequota-975-6663-crds' to 'e2e-test-e2e-resourcequota-975'"
      "context deadline exceeded" at 20:30:40 (~61s later)
      "Found 0 events" in namespace e2e-resourcequota-975
    - Build log: "failed: (1m1s) 2026-04-26T20:30:40 '[sig-api-machinery] ResourceQuota should create
      a ResourceQuota and capture the life of a custom resource.'"
    - Retry log: "Test [...] not eligible for retries (strategy returned 0)"

I'm rerunning

@sdminonne

Copy link
Copy Markdown
Contributor Author

/test e2e-conformance

@openshift-ci openshift-ci Bot added the needs-rebase Indicates a PR cannot be merged because it has merge conflicts with HEAD. label Apr 27, 2026
@enxebre

enxebre commented Apr 28, 2026

Copy link
Copy Markdown
Member

Thanks! few questions:

Add cluster-autoscaler.kubernetes.io/safe-to-evict: "true" annotation to the kas-connection-checker pod template so the cluster autoscaler can evict pods despite system-node-critical priority class.

why is this annotation needed? does the autoscaler has any specific behaviour for pods with this priority class during draining?

system-node-critical priority class, which prevents the cluster autoscaler from evicting its pods during scale-down.

where does autoscaler implement this?
wouldn't this mess with scale down draining just because of its hard tolerations?

@sdminonne
sdminonne force-pushed the fix/kas-connection-checker-pdb-safe-to-evict branch from 8979ea8 to a1b9129 Compare April 28, 2026 09:27
@openshift-ci openshift-ci Bot removed the needs-rebase Indicates a PR cannot be merged because it has merge conflicts with HEAD. label Apr 28, 2026
@sdminonne

Copy link
Copy Markdown
Contributor Author

/verified by @sdminonne in local dev cluster info added in comment

@openshift-ci-robot

Copy link
Copy Markdown

@sdminonne: This PR has been marked as verified by @sdminonne in local dev cluster info added in comment.

Details

In response to this:

/verified by @sdminonne in local dev cluster info added in comment

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@sdminonne

Copy link
Copy Markdown
Contributor Author

/retest

@sdminonne

Copy link
Copy Markdown
Contributor Author

/pipeline required

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Scheduling tests matching the pipeline_run_if_changed or not excluded by pipeline_skip_if_only_changed parameters:
/test e2e-aks-5-0
/test e2e-aws-5-0
/test e2e-aks
/test e2e-aws
/test e2e-aws-upgrade-hypershift-operator
/test e2e-kubevirt-aws-ovn-reduced
/test e2e-v2-aws
/test e2e-v2-azure-self-managed
/test e2e-v2-gke

@sdminonne

Copy link
Copy Markdown
Contributor Author

/test e2e-conformance

sdminonne and others added 5 commits September 3, 2026 08:17
…y spread constraint

- Add cluster-autoscaler.kubernetes.io/safe-to-evict: "true" annotation
  so the cluster autoscaler can evict these pods during scale-down.
- Add TopologySpreadConstraints to distribute replicas across nodes.
- Remove all custom tolerations to break the drain loop caused by the
  blanket NoSchedule toleration matching cordon taints.
- Rename reconcileKASConnectionCheckerDeployment to
  reconcileKASConnectionChecker and align test function name accordingly.
- Add EnsureKASConnectionCheckerSpec e2e helper gated behind
  CPOAtLeast(Version423) and wire it into TestCreateCluster.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Address review feedback: merge labels and annotations instead of
replacing the entire map, so annotations set by other operators are
not wiped out. Explicitly delete the stale openshift.io/required-scc
annotation from older HCCO versions (kube-system is exempt from SCC
admission via systemNamespaces allowlist). Rename test to follow
TESTING.md Test<FunctionName> convention.

Signed-off-by: Salvatore Dario Minonne <sminonne@redhat.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Move EnsureKASConnectionCheckerSpec from v1 create_cluster_test to the
v2 hosted_cluster_compliance_test using GetHostedClusterClient(). The
test validates safe-to-evict annotation, empty tolerations, and topology
spread constraint, skipping for CPO < 4.23.

Signed-off-by: Salvatore Dario Minonne <sminonne@redhat.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The reconcileKASConnectionChecker function manages a ServiceAccount,
ConfigMap, Deployment, Role, and RoleBinding. Update the log message
from "deployment" to "resources" to accurately reflect the scope.

Signed-off-by: Salvatore Dario Minonne <sminonne@redhat.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Salvatore Dario Minonne <sminonne@redhat.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@sdminonne

Copy link
Copy Markdown
Contributor Author

/retest

Replace GetHostedClusterClient with GetHostedClusterClientViaPortForward
in the kas-connection-checker spec compliance test. The port-forward
approach tunnels through the management cluster API to the kube-apiserver
pod, bypassing external DNS and load balancers. This makes the test work
for every provider and visibility mode, including after Azure endpoint
access transition tests cycle the topology.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@sdminonne

Copy link
Copy Markdown
Contributor Author

/test okd-scos-images

@sdminonne

Copy link
Copy Markdown
Contributor Author
/test e2e-v2-azure-self-managed

@sdminonne

Copy link
Copy Markdown
Contributor Author

/test okd-scos-images

@sdminonne

Copy link
Copy Markdown
Contributor Author

/test e2e-v2-azure-self-managed

@sdminonne

Copy link
Copy Markdown
Contributor Author

/pipeline required

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Scheduling tests matching the pipeline_run_if_changed or not excluded by pipeline_skip_if_only_changed parameters:
/test e2e-aks-5-0
/test e2e-aws-5-0
/test e2e-aks
/test e2e-aws
/test e2e-aws-upgrade-hypershift-operator
/test e2e-kubevirt-aws-ovn-reduced
/test e2e-v2-aws
/test e2e-v2-azure-self-managed
/test e2e-v2-gke

@sdminonne

Copy link
Copy Markdown
Contributor Author

/retest-required

@sdminonne

Copy link
Copy Markdown
Contributor Author

/test e2e-aks

@csrwng

csrwng commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

/lgtm

@sdminonne

Copy link
Copy Markdown
Contributor Author

/verified by @sdminonne in local dev cluster. Last commit touched only e2e

@openshift-ci-robot

Copy link
Copy Markdown

@sdminonne: This PR has been marked as verified by @sdminonne in local dev cluster. Last commit touched only e2e.

Details

In response to this:

/verified by @sdminonne in local dev cluster. Last commit touched only e2e

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-ci

openshift-ci Bot commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

@sdminonne: all tests passed!

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@openshift-ci-robot

Copy link
Copy Markdown

@sdminonne: Jira Issue Verification Checks: Jira Issue OCPBUGS-84368
✔️ This pull request was pre-merge verified.
✔️ All associated pull requests have merged.
✔️ All associated, merged pull requests were pre-merge verified.

Jira Issue OCPBUGS-84368 has been moved to the MODIFIED state and will move to the VERIFIED state when the change is available in an accepted nightly payload. 🕓

Details

In response to this:

Summary

Problem

The kas-connection-checker deployment runs in the kube-system namespace without a PDB or safe-to-evict annotation. The cluster autoscaler treats pods in kube-system without a PDB as system-critical pods that cannot be evicted during scale-down, blocking node draining on underutilized nodes. See also the autoscaler system drainability rule.

Additionally, the deployment used a blanket {Effect: NoSchedule, Operator: Exists} toleration that matched the cordon taint (node.kubernetes.io/unschedulable:NoSchedule), causing replacement pods to be scheduled back onto cordoned nodes during drain — creating a destructive evict-reschedule loop.

Solution

  1. Add cluster-autoscaler.kubernetes.io/safe-to-evict: "true" annotation to the pod template. This bypasses the autoscaler's system drainability rule for kube-system pods, allowing nodes to be scaled down. The annotation is preferred over a PDB because it doesn't block scale-to-zero operations.

  2. Remove all custom tolerations. The previous blanket NoSchedule toleration matched the cordon taint, causing the evict-reschedule loop described above. In a HyperShift hosted cluster there are no master nodes in the guest cluster, and the checker doesn't need to run on infra nodes, so no NoSchedule tolerations are needed. Kubernetes provides default NoExecute tolerations (300s grace period) for unreachable and not-ready via the DefaultTolerationSeconds admission controller.

  3. Add TopologySpreadConstraint with ScheduleAnyway to spread the 3 replicas across nodes. Using ScheduleAnyway (soft constraint) avoids blocking scheduling on small clusters (1-2 nodes).

  4. Preserve existing annotations and labels during reconciliation instead of replacing the entire map. Explicitly delete the stale openshift.io/required-scc annotation from older HCCO versions (kube-system is exempt from SCC admission via the systemNamespaces allowlist in origin).

Scope and impact

The tolerations are only removed from the kas-connection-checker Deployment in the guest cluster's kube-system namespace — no other workloads are affected. If all worker nodes carry custom NoSchedule taints and the checker pods cannot schedule, the ControlPlaneConnectionAvailable condition will report False (not Unknown) with reason ControlPlaneConnectionKASAccessFailedReason, because the ConfigMap never receives a lastSucceeded timestamp. This is an accurate signal: if the checker cannot run, data-plane-to-control-plane connectivity genuinely cannot be verified. In the normal case, the 3 replicas with ScheduleAnyway topology spread will schedule on any untainted nodes.

Changes

  • control-plane-operator/.../resources/resources.go: Add safe-to-evict annotation to pod template (merge, don't replace); remove stale required-scc annotation; remove all custom tolerations; add TopologySpreadConstraint; broaden log message from "deployment" to "resources".
  • control-plane-operator/.../resources/resources_test.go: Update tests to validate the safe-to-evict annotation, assert no custom tolerations, verify topology spread constraint. Rename test to TestReconcileKASConnectionChecker per TESTING.md convention.
  • test/e2e/v2/tests/hosted_cluster_compliance_test.go: Add EnsureKASConnectionCheckerSpecTest to the v2 e2e suite, gated behind VersionAtLeast(Version423).
  • test/e2e/create_cluster_test.go: Remove EnsureKASConnectionCheckerSpec call (moved to v2).
  • test/e2e/util/util.go: Remove EnsureKASConnectionCheckerSpec helper (moved to v2).

Test plan

  • go test ./control-plane-operator/hostedclusterconfigoperator/controllers/resources/... -count=1 passes
  • make verify passes
  • Verify on a live cluster that the annotation is present on kas-connection-checker pods
  • Verify cluster autoscaler can scale down nodes running kas-connection-checker pods

Fixes: https://redhat.atlassian.net/browse/OCPBUGS-84368

🤖 Generated with Claude Code

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-merge-robot

Copy link
Copy Markdown
Contributor

Fix included in release 5.1.0-0.nightly-2026-09-05-100309

@jparrill

jparrill commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

/cherry-pick release-5.0

@openshift-cherrypick-robot

Copy link
Copy Markdown

@jparrill: #8338 failed to apply on top of branch "release-5.0":

Applying: fix(kas-connection-checker): add safe-to-evict annotation and topology spread constraint
Using index info to reconstruct a base tree...
M	control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go
M	control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go
M	test/e2e/util/util.go
Falling back to patching base and 3-way merge...
Auto-merging control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go
CONFLICT (content): Merge conflict in control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources.go
Auto-merging control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go
CONFLICT (content): Merge conflict in control-plane-operator/hostedclusterconfigoperator/controllers/resources/resources_test.go
Auto-merging test/e2e/util/util.go
error: Failed to merge in the changes.
hint: Use 'git am --show-current-patch=diff' to see the failed patch
hint: When you have resolved this problem, run "git am --continue".
hint: If you prefer to skip this patch, run "git am --skip" instead.
hint: To restore the original branch and stop patching, run "git am --abort".
hint: Disable this message with "git config set advice.mergeConflict false"
Patch failed at 0001 fix(kas-connection-checker): add safe-to-evict annotation and topology spread constraint

Details

In response to this:

/cherry-pick release-5.0

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. area/control-plane-operator Indicates the PR includes changes for the control plane operator - in an OCP release area/hypershift-operator Indicates the PR includes changes for the hypershift operator and API - outside an OCP release area/testing Indicates the PR includes changes for e2e testing jira/severity-important Referenced Jira bug's severity is important for the branch this PR is targeting. jira/valid-bug Indicates that a referenced Jira bug is valid for the branch this PR is targeting. jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. lgtm Indicates that a PR is ready to be merged. verified Signifies that the PR passed pre-merge verification criteria

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants