Skip to content

CNTRLPLANE-3276: Add Azure ExternalPrivateService and endpoint access transition test - #8718

Merged
openshift-merge-bot[bot] merged 2 commits into
openshift:mainfrom
bryan-cox:CNTRLPLANE-3276
Jun 24, 2026
Merged

openshift-merge-bot[bot] merged 2 commits into
openshift:mainfrom
bryan-cox:CNTRLPLANE-3276

Conversation

@bryan-cox

@bryan-cox bryan-cox commented Jun 11, 2026 •

Copy link
Copy Markdown
Member

What this PR does / why we need it:

Adds ExternalPrivateService support to the Azure PLS controller and a v2 e2e lifecycle test validating topology transitions on Azure self-managed clusters.

Controller changes (control-plane-operator)

The Azure PLS controller now creates ExternalName Services with external-dns.alpha.kubernetes.io/hostname annotations when a cluster is Private with custom Route hostnames configured. This brings Azure to parity with AWS and GCP private link controllers:

  • Uses PE IP as ExternalName target (matching GCP PSC pattern, not CNAME like AWS)
  • Labels services with hypershift.openshift.io/azure-pls-external-private-svc for lifecycle management
  • Deletes services by label when cluster transitions to public (IsPublicHCP)
  • Sets owner reference to HCP for garbage collection
  • Only reconciles for the private-router CR

Unit tests cover all three new functions: hcpExternalNames, reconcileExternalService, and reconcileExternalServices.

E2E test (test/e2e/v2)

Validates transitioning a HostedCluster between Private and PublicAndPrivate topology:

  1. Private state: Verifies KAS and OAuth ExternalPrivateService exist (ExternalName type)
  2. Private → PublicAndPrivate: Confirms public routes created, private routes deleted, ExternalPrivateService cleaned up, PLS CRs persist, API reachable via public route
  3. PublicAndPrivate → Private: Confirms private routes recreated, ExternalPrivateService recreated, PLS CRs persist, cluster Available and not Degraded

DeferCleanup restores Private topology if the test fails mid-transition.

Which issue(s) this PR fixes:

Fixes https://issues.redhat.com/browse/CNTRLPLANE-3276

Special notes for your reviewer:

  • ExternalPrivateService implementation mirrors the established AWS/GCP pattern
  • Uses controllerutil.CreateOrUpdate (Azure PLS controller doesn't embed the upsert provider)
  • ExternalPrivateService convergence after restore to Private takes ~3 minutes (controller reconcile cycle)
  • No changes to lifecycle/azure.go — the private variant's LabelFilter already includes self-managed-azure-private
  • DeferCleanup uses context.Background() intentionally — cleanup may run after testCtx.Context is cancelled on timeout

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Pipeline controller notification
This repo is configured to use the pipeline controller. Second-stage tests will be triggered either automatically or after lgtm label is added, depending on the repository configuration. The pipeline controller will automatically detect which contexts are required and will utilize /test Prow commands to trigger the second stage.

For optional jobs, comment /test ? to see a list of all defined jobs. To trigger manually all jobs from second stage use /pipeline required command.

This repository is configured in: LGTM mode

@openshift-ci-robot openshift-ci-robot added the jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. label Jun 11, 2026
@openshift-ci-robot

openshift-ci-robot commented Jun 11, 2026 •

Copy link
Copy Markdown

@bryan-cox: This pull request references CNTRLPLANE-3276 which is a valid jira issue.

Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the task to target the "5.0.0" version, but no target version was set.

Details

In response to this:

What this PR does / why we need it:

Adds a v2 e2e lifecycle test that validates transitioning a HostedCluster between Private and PublicAndPrivate topology on Azure self-managed clusters. The test verifies:

  1. Private → PublicAndPrivate: Updates topology, waits for PublicEndpointExposed condition to become True with reason SharedIngressConfigured, confirms KAS service becomes LoadBalancer, validates PLS CRs persist, and checks API reachability.
  2. PublicAndPrivate → Private: Restores topology, waits for PublicEndpointExposed condition to become False with reason TopologyPrivate, confirms KAS service returns to ClusterIP, validates PLS CRs persist.
  3. Private connectivity verification: Confirms API server is reachable via the private path after restore.

The test runs on the existing private cluster variant (labeled self-managed-azure-private) after AzurePrivateTopologyTest completes. DeferCleanup restores Private topology if the test fails mid-transition.

Which issue(s) this PR fixes:

Fixes https://issues.redhat.com/browse/CNTRLPLANE-3276

Special notes for your reviewer:

  • No changes to lifecycle/azure.go — the private variant's LabelFilter already includes self-managed-azure-private
  • Uses e2eutil.ConditionPredicate[*hyperv1.HostedCluster] rather than a custom condition helper
  • DeferCleanup uses context.Background() intentionally — cleanup may run after testCtx.Context is cancelled on timeout
  • The sync.Once-cached guest client in GetHostedClusterClient() is safe across topology transitions because the private endpoint stays active in PublicAndPrivate mode

Checklist:

  • Subject and description added to both, commit and PR.
  • Relevant issues have been referenced.
  • This change includes docs.
  • This change includes unit tests.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-ci openshift-ci Bot added the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Jun 11, 2026
@openshift-ci

openshift-ci Bot commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

Skipping CI for Draft Pull Request.
If you want CI signal for your change, please convert it to an actual PR.
You can still manually trigger a test run with /test all

@coderabbitai

coderabbitai Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

This PR adds a new e2e test, AzureEndpointAccessTransitionTest, that validates HostedCluster topology transitions on Azure. The test transitions a cluster from Private to PublicAndPrivate and back to Private, verifying that KAS external routes appear and are deleted appropriately for each topology state, AzurePrivateLinkService CRs persist with non-empty Status.PrivateLinkServiceAlias, and the hosted cluster API remains reachable throughout. Supporting helper predicates match service types, verify PLS alias presence, and poll API reachability via namespace listing. The test is registered immediately after AzurePrivateTopologyTest.

Suggested reviewers

  • cblecker
  • Nirshal
  • csrwng
🚥 Pre-merge checks | ✅ 11
✅ Passed checks (11 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Stable And Deterministic Test Names ✅ Passed All test names in AzureEndpointAccessTransitionTest are static descriptive strings without dynamic values, pod names, timestamps, UUIDs, node names, or IP addresses.
Test Structure And Quality ✅ Passed Test code follows all quality requirements: single responsibility per test, proper setup/cleanup via BeforeAll/DeferCleanup, timeouts on all cluster operations, meaningful assertion messages with f...
Topology-Aware Scheduling Compatibility ✅ Passed This PR adds only e2e test code (test/e2e/v2/tests/hosted_cluster_azure_test.go) with no deployment manifests, operator code, or controllers. The custom check's scope explicitly applies to these, m...
Ipv6 And Disconnected Network Test Compatibility ✅ Passed Test contains no hardcoded IPv4 addresses, IPv4-specific logic, or external connectivity requirements. All operations are cluster-internal Kubernetes API calls.
No-Weak-Crypto ✅ Passed The PR adds an e2e test for Azure topology transitions with no weak cryptography usage. No MD5, SHA1, DES, RC4, 3DES, Blowfish, ECB algorithms, custom crypto implementations, or non-constant-time s...
Container-Privileges ✅ Passed PR adds a Go e2e test file with no container/K8s manifests or privilege configurations. Check is not applicable to test code.
No-Sensitive-Data-In-Logs ✅ Passed Logging statements in AzureEndpointAccessTransitionTest log only non-sensitive data: Route hostnames (DNS names) and Azure PLS aliases (service metadata). No passwords, tokens, API keys, or PII are...
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: adding an Azure endpoint access transition test for ExternalPrivateService functionality.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@openshift-ci

openshift-ci Bot commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: bryan-cox

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@openshift-ci openshift-ci Bot added approved Indicates a PR has been approved by an approver from all required OWNERS files. area/platform/azure PR/issue for Azure (AzurePlatform) platform area/testing Indicates the PR includes changes for e2e testing and removed do-not-merge/needs-area labels Jun 11, 2026
@bryan-cox

Copy link
Copy Markdown
Member Author

/test e2e-azure-v2-self-managed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
test/e2e/v2/tests/hosted_cluster_azure_test.go (1)

270-280: 💤 Low value

Add WithInterval for polling consistency.

The condition check at lines 253-268 specifies WithInterval(15*time.Second), but this EventuallyObject call omits it. For consistent polling behavior across similar checks in this test, consider adding the interval.

Suggested fix
 			e2eutil.EventuallyObject(GinkgoTB(), ctx, "KAS service is LoadBalancer",
 				func(ctx context.Context) (*corev1.Service, error) {
 					svc := hcpmanifests.KubeAPIServerService(controlPlaneNamespace)
 					err := testCtx.MgmtClient.Get(ctx, crclient.ObjectKeyFromObject(svc), svc)
 					return svc, err
 				},
 				[]e2eutil.Predicate[*corev1.Service]{
 					serviceTypePredicate(corev1.ServiceTypeLoadBalancer),
 				},
 				e2eutil.WithTimeout(10*time.Minute),
+				e2eutil.WithInterval(15*time.Second),
 			)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/e2e/v2/tests/hosted_cluster_azure_test.go` around lines 270 - 280, Add a
consistent polling interval to the EventuallyObject call by including
e2eutil.WithInterval(15*time.Second) alongside the existing e2eutil.WithTimeout
option; locate the call to e2eutil.EventuallyObject (the block returning
*corev1.Service and using serviceTypePredicate) and append the WithInterval
option to its variadic options so it uses the same 15s polling cadence as the
earlier check.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@test/e2e/v2/tests/hosted_cluster_azure_test.go`:
- Around line 270-280: Add a consistent polling interval to the EventuallyObject
call by including e2eutil.WithInterval(15*time.Second) alongside the existing
e2eutil.WithTimeout option; locate the call to e2eutil.EventuallyObject (the
block returning *corev1.Service and using serviceTypePredicate) and append the
WithInterval option to its variadic options so it uses the same 15s polling
cadence as the earlier check.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 3b2222e8-a730-4864-a308-b08be14365df

📥 Commits

Reviewing files that changed from the base of the PR and between 35c0190 and 8a9c92f.

📒 Files selected for processing (1)
  • test/e2e/v2/tests/hosted_cluster_azure_test.go

@codecov

codecov Bot commented Jun 11, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.16129% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 41.91%. Comparing base (28adcd6) to head (8efac9c).
⚠️ Report is 31 commits behind head on main.

Files with missing lines Patch % Lines
.../controllers/azureprivatelinkservice/controller.go 95.16% 2 Missing and 1 partial ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #8718      +/-   ##
==========================================
- Coverage   42.18%   41.91%   -0.28%     
==========================================
  Files         767      769       +2     
  Lines       95092    96763    +1671     
==========================================
+ Hits        40114    40557     +443     
- Misses      52163    53402    +1239     
+ Partials     2815     2804      -11     
Files with missing lines Coverage Δ
.../controllers/azureprivatelinkservice/controller.go 92.38% <95.16%> (+0.25%) ⬆️

... and 7 files with indirect coverage changes

Flag Coverage Δ
cmd-support 35.46% <ø> (+<0.01%) ⬆️
cpo-hostedcontrolplane 44.84% <ø> (ø)
cpo-other 44.56% <95.16%> (+0.31%) ⬆️
hypershift-operator 50.17% <ø> (-1.75%) ⬇️
other 31.69% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@bryan-cox

Copy link
Copy Markdown
Member Author

/test e2e-azure-v2-self-managed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
test/e2e/v2/tests/hosted_cluster_azure_test.go (2)

279-309: ⚡ Quick win

Consider verifying HostedCluster condition PublicEndpointExposed returns to False.

The PR description states the test should "wait for PublicEndpointExposed to become False with reason TopologyPrivate" when transitioning back to Private. Currently, the test only verifies the KAS service type returns to ClusterIP. Adding a condition check would ensure the full reconciliation completed.

📋 Suggested addition to verify condition

After line 297, add:

 		e2eutil.EventuallyObject(GinkgoTB(), ctx, "KAS service is ClusterIP",
 			func(ctx context.Context) (*corev1.Service, error) {
 				svc := hcpmanifests.KubeAPIServerService(controlPlaneNamespace)
 				err := testCtx.MgmtClient.Get(ctx, crclient.ObjectKeyFromObject(svc), svc)
 				return svc, err
 			},
 			[]e2eutil.Predicate[*corev1.Service]{
 				serviceTypePredicate(corev1.ServiceTypeClusterIP),
 			},
 			e2eutil.WithTimeout(10*time.Minute),
 		)
+
+		e2eutil.EventuallyObject(GinkgoTB(), ctx, "HostedCluster PublicEndpointExposed condition is False",
+			func(ctx context.Context) (*hyperv1.HostedCluster, error) {
+				latest := &hyperv1.HostedCluster{}
+				err := testCtx.MgmtClient.Get(ctx, crclient.ObjectKeyFromObject(hc), latest)
+				return latest, err
+			},
+			[]e2eutil.Predicate[*hyperv1.HostedCluster]{
+				e2eutil.ConditionPredicate[*hyperv1.HostedCluster](
+					hyperv1.PublicEndpointExposed,
+					metav1.ConditionFalse,
+					hyperv1.TopologyPrivateReason,
+				),
+			},
+			e2eutil.WithTimeout(10*time.Minute),
+		)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/e2e/v2/tests/hosted_cluster_azure_test.go` around lines 279 - 309, Add
an assertion that the HostedCluster's PublicEndpointExposed condition flips to
False with reason TopologyPrivate after topology revert: use
e2eutil.EventuallyObject (like the existing KAS service check) to fetch the
HostedCluster (variable hc via testCtx.MgmtClient) and poll until
conditions.Get(hc.Status.Conditions, hyperv1.PublicEndpointExposed).Status ==
corev1.ConditionFalse and .Reason == "TopologyPrivate" (or use the helper that
checks condition values), with an appropriate timeout placed after the KAS
service ClusterIP check to ensure full reconciliation completed.

245-277: ⚡ Quick win

Consider verifying HostedCluster condition PublicEndpointExposed.

The PR description states the test should "wait for HostedCluster condition PublicEndpointExposed to become True with reason SharedIngressConfigured." Currently, the test only verifies the KAS service type changes to LoadBalancer. Adding a condition check would provide stronger validation that the control plane reconciliation loop completed successfully.

📋 Suggested addition to verify condition

After line 263, add a condition check using e2eutil.ConditionPredicate:

 		e2eutil.EventuallyObject(GinkgoTB(), ctx, "KAS service is LoadBalancer",
 			func(ctx context.Context) (*corev1.Service, error) {
 				svc := hcpmanifests.KubeAPIServerService(controlPlaneNamespace)
 				err := testCtx.MgmtClient.Get(ctx, crclient.ObjectKeyFromObject(svc), svc)
 				return svc, err
 			},
 			[]e2eutil.Predicate[*corev1.Service]{
 				serviceTypePredicate(corev1.ServiceTypeLoadBalancer),
 			},
 			e2eutil.WithTimeout(10*time.Minute),
 		)
+
+		e2eutil.EventuallyObject(GinkgoTB(), ctx, "HostedCluster PublicEndpointExposed condition is True",
+			func(ctx context.Context) (*hyperv1.HostedCluster, error) {
+				latest := &hyperv1.HostedCluster{}
+				err := testCtx.MgmtClient.Get(ctx, crclient.ObjectKeyFromObject(hc), latest)
+				return latest, err
+			},
+			[]e2eutil.Predicate[*hyperv1.HostedCluster]{
+				e2eutil.ConditionPredicate[*hyperv1.HostedCluster](
+					hyperv1.PublicEndpointExposed,
+					metav1.ConditionTrue,
+					hyperv1.SharedIngressConfiguredReason,
+				),
+			},
+			e2eutil.WithTimeout(10*time.Minute),
+		)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/e2e/v2/tests/hosted_cluster_azure_test.go` around lines 245 - 277, Add a
check that waits for the HostedCluster condition "PublicEndpointExposed" to be
True with reason "SharedIngressConfigured" (using the existing HostedCluster
object hc and testCtx.MgmtClient) to ensure control-plane reconciliation
completed; implement this by calling e2eutil.EventuallyObject (or the existing
helper that waits on conditions) with e2eutil.ConditionPredicate for condition
"PublicEndpointExposed" and reason "SharedIngressConfigured" (use
e2eutil.ConditionPredicate(hc, "PublicEndpointExposed", corev1.ConditionTrue,
"SharedIngressConfigured") or the appropriate signature) placed after the KAS
service LoadBalancer check/PLS verification and before verifyAPIReachable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@test/e2e/v2/tests/hosted_cluster_azure_test.go`:
- Around line 279-309: Add an assertion that the HostedCluster's
PublicEndpointExposed condition flips to False with reason TopologyPrivate after
topology revert: use e2eutil.EventuallyObject (like the existing KAS service
check) to fetch the HostedCluster (variable hc via testCtx.MgmtClient) and poll
until conditions.Get(hc.Status.Conditions, hyperv1.PublicEndpointExposed).Status
== corev1.ConditionFalse and .Reason == "TopologyPrivate" (or use the helper
that checks condition values), with an appropriate timeout placed after the KAS
service ClusterIP check to ensure full reconciliation completed.
- Around line 245-277: Add a check that waits for the HostedCluster condition
"PublicEndpointExposed" to be True with reason "SharedIngressConfigured" (using
the existing HostedCluster object hc and testCtx.MgmtClient) to ensure
control-plane reconciliation completed; implement this by calling
e2eutil.EventuallyObject (or the existing helper that waits on conditions) with
e2eutil.ConditionPredicate for condition "PublicEndpointExposed" and reason
"SharedIngressConfigured" (use e2eutil.ConditionPredicate(hc,
"PublicEndpointExposed", corev1.ConditionTrue, "SharedIngressConfigured") or the
appropriate signature) placed after the KAS service LoadBalancer check/PLS
verification and before verifyAPIReachable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 9e7d7d98-d5d6-4677-8778-67d0d9a0b9a5

📥 Commits

Reviewing files that changed from the base of the PR and between 8a9c92f and 13e0579.

📒 Files selected for processing (1)
  • test/e2e/v2/tests/hosted_cluster_azure_test.go

@bryan-cox

Copy link
Copy Markdown
Member Author

/test e2e-azure-v2-self-managed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
test/e2e/v2/tests/hosted_cluster_azure_test.go (1)

246-314: ⚡ Quick win

Add labels to each new It for filterability consistency.

The new specs are under a labeled Context, but the It blocks themselves are unlabeled.

Proposed fix
-		It("should transition from Private to PublicAndPrivate", func() {
+		It("should transition from Private to PublicAndPrivate", Label("azure-endpoint-transition"), func() {
...
-		It("should transition from PublicAndPrivate back to Private", func() {
+		It("should transition from PublicAndPrivate back to Private", Label("azure-endpoint-transition"), func() {
...
-		It("should reach the API server after restoring Private topology", func() {
+		It("should reach the API server after restoring Private topology", Label("azure-endpoint-transition"), func() {

As per coding guidelines, apply labels to both Describe and It blocks for test filtering.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/e2e/v2/tests/hosted_cluster_azure_test.go` around lines 246 - 314, Add
labels to the three It blocks in the test to maintain filtering consistency with
the parent Context block. The It blocks "should transition from Private to
PublicAndPrivate", "should transition from PublicAndPrivate back to Private",
and "should reach the API server after restoring Private topology" need to be
labeled using the Label() function according to coding guidelines for test
filtering. Apply appropriate labels that categorize these topology transition
tests consistently.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@test/e2e/v2/tests/hosted_cluster_azure_test.go`:
- Around line 295-297: The assertion for the service type check is too loose and
only verifies that the type is not LoadBalancer, allowing false positives from
other unintended types like NodePort or ExternalName. Replace the
NotTo(Equal(corev1.ServiceTypeLoadBalancer)) check with explicit assertions that
verify the service is in one of the expected states after restore: either the
service should not be found (NotFound error) or it should have type ClusterIP.
This ensures the assertion catches actual regressions instead of passing for any
non-LoadBalancer type.
- Around line 236-243: The DeferCleanup function currently logs a warning when
the topology restore fails (when restoreErr is not nil) but continues execution,
which can leave the cluster state mutated and cause cascading failures in
subsequent tests. Modify the error handling block where restoreErr is checked to
fail the test using an appropriate Ginkgo assertion or failure method (such as
GinkgoTB().Fail) instead of just logging a warning, ensuring that any failure to
restore the Azure topology causes the test cleanup to fail and prevent state
corruption from affecting other specs.

---

Nitpick comments:
In `@test/e2e/v2/tests/hosted_cluster_azure_test.go`:
- Around line 246-314: Add labels to the three It blocks in the test to maintain
filtering consistency with the parent Context block. The It blocks "should
transition from Private to PublicAndPrivate", "should transition from
PublicAndPrivate back to Private", and "should reach the API server after
restoring Private topology" need to be labeled using the Label() function
according to coding guidelines for test filtering. Apply appropriate labels that
categorize these topology transition tests consistently.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 82fff169-72ca-4e38-910e-d973910bb5f5

📥 Commits

Reviewing files that changed from the base of the PR and between 13e0579 and 57c8e8e.

📒 Files selected for processing (1)
  • test/e2e/v2/tests/hosted_cluster_azure_test.go

Comment thread test/e2e/v2/tests/hosted_cluster_azure_test.go
Comment thread test/e2e/v2/tests/hosted_cluster_azure_test.go Outdated
@bryan-cox

Copy link
Copy Markdown
Member Author

/test e2e-azure-v2-self-managed

1 similar comment
@bryan-cox

Copy link
Copy Markdown
Member Author

/test e2e-azure-v2-self-managed

@bryan-cox

Copy link
Copy Markdown
Member Author

/test pull-ci-openshift-hypershift-main-e2e-azure-v2-self-managed

@bryan-cox

Copy link
Copy Markdown
Member Author

/test e2e-azure-v2-self-managed

@bryan-cox

Copy link
Copy Markdown
Member Author

/test e2e-azure-v2-self-managed

@hypershift-jira-solve-ci

Copy link
Copy Markdown
Contributor

AI Test Failure Analysis

Job: pull-ci-openshift-hypershift-main-e2e-aks | Build: 2069433019269124096 | Cost: $3.5601244999999997 | Failed step: hypershift-azure-run-e2e

View full analysis report


Generated by hypershift-analyze-e2e-failure post-step using Claude claude-opus-4-6

@bryan-cox

Copy link
Copy Markdown
Member Author

/retest

@bryan-cox

Copy link
Copy Markdown
Member Author

/verified by e2e

@openshift-ci-robot openshift-ci-robot added the verified Signifies that the PR passed pre-merge verification criteria label Jun 23, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@bryan-cox: This PR has been marked as verified by e2e.

Details

In response to this:

/verified by e2e

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@bryan-cox

Copy link
Copy Markdown
Member Author

/retest

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

/retest-required

Remaining retests: 0 against base HEAD 9a16fd2 and 2 for PR HEAD 8efac9c in total

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

/retest-required

Remaining retests: 0 against base HEAD 49cdfcf and 1 for PR HEAD 8efac9c in total

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

/retest-required

Remaining retests: 0 against base HEAD bc3bda9 and 0 for PR HEAD 8efac9c in total

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

/hold

Revision 8efac9c was retested 3 times: holding

@openshift-ci openshift-ci Bot added the do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. label Jun 24, 2026
@bryan-cox

Copy link
Copy Markdown
Member Author

/test e2e-aws

@bryan-cox

Copy link
Copy Markdown
Member Author

/retest

ci infra issue before

@bryan-cox

Copy link
Copy Markdown
Member Author

/override codecov/project

@openshift-ci

openshift-ci Bot commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

@bryan-cox: Overrode contexts on behalf of bryan-cox: codecov/project

Details

In response to this:

/override codecov/project

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@bryan-cox

Copy link
Copy Markdown
Member Author

/hold cancel

@openshift-ci openshift-ci Bot removed the do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. label Jun 24, 2026
@hypershift-jira-solve-ci

hypershift-jira-solve-ci Bot commented Jun 24, 2026 •

Copy link
Copy Markdown
Contributor

Test Failure Analysis Complete

Job Information

  • Prow Job: codecov/project
  • Build ID: 83221405394
  • PR: #8718 — CNTRLPLANE-3276: Add Azure ExternalPrivateService and endpoint access transition test
  • Check: codecov/project (Codecov App)
  • Head Commit: 8efac9c
  • Base Commit: 28adcd6

Test Failure Analysis

Error

codecov/project: 41.91% (-0.28%) compared to 28adcd6
Patch coverage is 95.16129% with 3 lines in your changes missing coverage.
⚠ Report is 31 commits behind head on main.

Summary

This is not a test failure — it is a Codecov project-level coverage gate failure caused by a stale base comparison. Codecov's report is 31 commits behind main, meaning it compared the PR against an outdated snapshot. The 31 unrelated commits merged to main added more uncovered lines (+1,239 misses) than covered (+443 hits), dragging overall project coverage from 42.18% to 41.91% (-0.28%). The PR's own patch coverage is a healthy 95.16% — only 3 lines out of ~62 in controller.go lack coverage. A Prow override has already marked a second codecov/project check as success, so this failure is non-blocking.

Root Cause

The root cause is a stale Codecov base reference, not a deficiency in the PR's test coverage.

  1. Stale base comparison (primary cause): Codecov's report explicitly warns it is "31 commits behind head on main." Codecov compared this PR against a snapshot of main from 31 commits ago (28adcd6). The 31 commits merged to main between that snapshot and the current HEAD introduced +1,671 total lines across 769 files. Only ~62 lines come from this PR's patch — the remaining ~1,609 lines are from unrelated commits that had a net effect of adding more uncovered lines (+1,239 misses) than covered lines (+443 hits), dragging overall project coverage from 42.18% down to 41.91%.

  2. Codecov default threshold: The repo's Codecov configuration does not define an explicit coverage.status.project.default.threshold value. Codecov's default behavior is to fail the project check when coverage decreases by any amount — even a 0.01% drop triggers failure.

  3. PR patch coverage is healthy: The PR's own patch achieves 95.16% coverage. The 3 uncovered lines are in control-plane-operator/controllers/azureprivatelinkservice/controller.go — likely nil-guard initialization branches (e.g., if svc.Labels == nil / if svc.Annotations == nil) that are low-value to test explicitly since the fake client builder initializes maps.

  4. Already overridden: A Prow override (/override codecov/project by bryan-cox) has already set a successful codecov/project status, making this failure non-blocking for merge.

Recommendations
  1. No action required on this PR — The failure is caused by Codecov's stale base, not by insufficient test coverage. The override is already in place and the PR's 95.16% patch coverage exceeds any reasonable threshold.

  2. Optional: Add a coverage threshold to codecov.yml — To prevent future false failures from stale base comparisons, the repo could add an explicit threshold tolerance:

    coverage:
      status:
        project:
          default:
            threshold: 0.5%  # Allow up to 0.5% drop
  3. Optional: Cover the 3 remaining lines — The 2 missing + 1 partial line in controller.go are likely the nil-map guard branches which are low-value to test explicitly.

  4. Merge when ready — The override is in place, and the codecov/project failure does not indicate any real quality issue with this PR.

Evidence
Evidence Detail
Codecov conclusion failure — 41.91% (-0.28%) compared to base 28adcd6
Stale base warning "Report is 31 commits behind head on main"
Total line delta +1,671 lines (PR contributes ~62, rest from 31 unrelated commits)
Coverage stats Hits: +443, Misses: +1,239, Partials: -11
PR patch coverage 95.16% — 3 lines missing in controller.go
File with gaps control-plane-operator/controllers/azureprivatelinkservice/controller.go (2 missing, 1 partial)
Prow override /override codecov/project by bryan-cox → success
Changed files controller.go, controller_test.go, hosted_cluster_azure_test.go

@openshift-ci

openshift-ci Bot commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

@bryan-cox: all tests passed!

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@openshift-merge-bot
openshift-merge-bot Bot merged commit 438c61f into openshift:main Jun 24, 2026
43 of 44 checks passed
@bryan-cox
bryan-cox deleted the CNTRLPLANE-3276 branch June 24, 2026 15:02
Nirshal added a commit to Nirshal/hypershift that referenced this pull request Jul 20, 2026
The private cluster variant had --oauth-publishing-strategy=LoadBalancer
set at creation time, which conflicts with AzureEndpointAccessTransitionTest
(PR openshift#8718) that expects Route-based OAuth on the private cluster. Since
ServicePublishingStrategy is immutable after cluster creation, the only
solution is a dedicated cluster variant.

- Remove --oauth-publishing-strategy=LoadBalancer from "private" ClusterSpec
- Add new "oauth-lb-private" ClusterSpec combining Private endpoint access
  with LoadBalancer OAuth publishing
- Move "private" and "oauth-lb-private" TestGroups from Sequential to
  Parallel since they now run on separate clusters
- Remove "private-and-oauth" SequentialGroup (no longer needed)

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
Commit-Message-Assisted-by: Claude (via Claude Code)
Nirshal added a commit to Nirshal/hypershift that referenced this pull request Jul 23, 2026
The private cluster variant had --oauth-publishing-strategy=LoadBalancer
set at creation time, which conflicts with AzureEndpointAccessTransitionTest
(PR openshift#8718) that expects Route-based OAuth on the private cluster. Since
ServicePublishingStrategy is immutable after cluster creation, the only
solution is a dedicated cluster variant.

- Remove --oauth-publishing-strategy=LoadBalancer from "private" ClusterSpec
- Add new "oauth-lb-private" ClusterSpec combining Private endpoint access
  with LoadBalancer OAuth publishing
- Move "private" and "oauth-lb-private" TestGroups from Sequential to
  Parallel since they now run on separate clusters
- Remove "private-and-oauth" SequentialGroup (no longer needed)

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
Commit-Message-Assisted-by: Claude (via Claude Code)
Nirshal added a commit to Nirshal/hypershift that referenced this pull request Jul 27, 2026
The private cluster variant had --oauth-publishing-strategy=LoadBalancer
set at creation time, which conflicts with AzureEndpointAccessTransitionTest
(PR openshift#8718) that expects Route-based OAuth on the private cluster. Since
ServicePublishingStrategy is immutable after cluster creation, the only
solution is a dedicated cluster variant.

- Remove --oauth-publishing-strategy=LoadBalancer from "private" ClusterSpec
- Add new "oauth-lb-private" ClusterSpec combining Private endpoint access
  with LoadBalancer OAuth publishing
- Move "private" and "oauth-lb-private" TestGroups from Sequential to
  Parallel since they now run on separate clusters
- Remove "private-and-oauth" SequentialGroup (no longer needed)

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
Commit-Message-Assisted-by: Claude (via Claude Code)
Nirshal added a commit to Nirshal/hypershift that referenced this pull request Jul 31, 2026
The private cluster variant had --oauth-publishing-strategy=LoadBalancer
set at creation time, which conflicts with AzureEndpointAccessTransitionTest
(PR openshift#8718) that expects Route-based OAuth on the private cluster. Since
ServicePublishingStrategy is immutable after cluster creation, the only
solution is a dedicated cluster variant.

- Remove --oauth-publishing-strategy=LoadBalancer from "private" ClusterSpec
- Add new "oauth-lb-private" ClusterSpec combining Private endpoint access
  with LoadBalancer OAuth publishing
- Move "private" and "oauth-lb-private" TestGroups from Sequential to
  Parallel since they now run on separate clusters
- Remove "private-and-oauth" SequentialGroup (no longer needed)

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
Commit-Message-Assisted-by: Claude (via Claude Code)
Nirshal added a commit to Nirshal/hypershift that referenced this pull request Aug 3, 2026
The private cluster variant had --oauth-publishing-strategy=LoadBalancer
set at creation time, which conflicts with AzureEndpointAccessTransitionTest
(PR openshift#8718) that expects Route-based OAuth on the private cluster. Since
ServicePublishingStrategy is immutable after cluster creation, the only
solution is a dedicated cluster variant.

- Remove --oauth-publishing-strategy=LoadBalancer from "private" ClusterSpec
- Add new "oauth-lb-private" ClusterSpec combining Private endpoint access
  with LoadBalancer OAuth publishing
- Move "private" and "oauth-lb-private" TestGroups from Sequential to
  Parallel since they now run on separate clusters
- Remove "private-and-oauth" SequentialGroup (no longer needed)

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
Commit-Message-Assisted-by: Claude (via Claude Code)
Nirshal added a commit to Nirshal/hypershift that referenced this pull request Aug 3, 2026
The private cluster variant had --oauth-publishing-strategy=LoadBalancer
set at creation time, which conflicts with AzureEndpointAccessTransitionTest
(PR openshift#8718) that expects Route-based OAuth on the private cluster. Since
ServicePublishingStrategy is immutable after cluster creation, the only
solution is a dedicated cluster variant.

- Remove --oauth-publishing-strategy=LoadBalancer from "private" ClusterSpec
- Add new "oauth-lb-private" ClusterSpec combining Private endpoint access
  with LoadBalancer OAuth publishing
- Move "private" and "oauth-lb-private" TestGroups from Sequential to
  Parallel since they now run on separate clusters
- Remove "private-and-oauth" SequentialGroup (no longer needed)

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
Commit-Message-Assisted-by: Claude (via Claude Code)
Nirshal added a commit to Nirshal/hypershift that referenced this pull request Aug 5, 2026
The private cluster variant had --oauth-publishing-strategy=LoadBalancer
set at creation time, which conflicts with AzureEndpointAccessTransitionTest
(PR openshift#8718) that expects Route-based OAuth on the private cluster. Since
ServicePublishingStrategy is immutable after cluster creation, the only
solution is a dedicated cluster variant.

- Remove --oauth-publishing-strategy=LoadBalancer from "private" ClusterSpec
- Add new "oauth-lb-private" ClusterSpec combining Private endpoint access
  with LoadBalancer OAuth publishing
- Move "private" and "oauth-lb-private" TestGroups from Sequential to
  Parallel since they now run on separate clusters
- Remove "private-and-oauth" SequentialGroup (no longer needed)

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
Commit-Message-Assisted-by: Claude (via Claude Code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. area/control-plane-operator Indicates the PR includes changes for the control plane operator - in an OCP release area/platform/azure PR/issue for Azure (AzurePlatform) platform area/testing Indicates the PR includes changes for e2e testing jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. lgtm Indicates that a PR is ready to be merged. verified Signifies that the PR passed pre-merge verification criteria

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants