Skip to content

OCPBUGS-120704: [release-4.22] fix(operator): health probes + pki-operator TLS config - #9600

Merged
openshift-merge-bot[bot] merged 3 commits into
openshift:release-4.22from
jparrill:ci-HealthProbeBindAddress
Sep 23, 2026
Merged

openshift-merge-bot[bot] merged 3 commits into
openshift:release-4.22from
jparrill:ci-HealthProbeBindAddress

Conversation

@jparrill

@jparrill jparrill commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Three cherry-picks to release-4.22 that fix both e2e-aws-upgrade-hypershift-operator and e2e-v2-aws presubmits (both 0/10 BROKEN, blocking the merge pipeline).

Both functional fixes are needed in the same PR to avoid a chicken-and-egg deadlock: neither presubmit can pass without both changes present, so merging them separately is impossible.


Commit 1: fix(operator): use dedicated health probes

Cherry-pick of 4084a520609 (original PR #9387, backported to 5.0 via #9534).

Root cause: The upgrade test binary (hypershift-tests:latest) is built from main, not from the release branch. Main's install code (cmd/install/assets/hypershift_operator.go) changed HO deployment probes from /metrics:9000 to /healthz:8081 and /readyz:8081. When the upgrade test deploys the new HO, it sets port 8081 probes, but the release-4.22 operator binary lacked HealthProbeBindAddress — no health server on 8081 — readiness probe fails — rollout stuck at 1/2 replicas — timeout after 608s.

Failure chain (empirically verified):

  1. CI runs e2e-aws-upgrade-hypershift-operator on release-4.22 PR
  2. Initial HO install uses 4.22's hcp CLI (INSTALL_FROM_LATEST=true) → probes at /metrics:9000 → works
  3. Upgrade step uses hypershift-tests:latest (from main) → deploys HO with /healthz:8081 and /readyz:8081 probes
  4. Release-4.22 operator binary has no HealthProbeBindAddress → port 8081 never opens
  5. Readiness probe fails → new replica never becomes ready → rollout stuck at 1/2
  6. Test timeout at 608s → FAIL

Changes (adapted to 4.22's monolithic run() structure):

  • hypershift-operator/main.go: add HealthProbeBindAddress, setupHealthChecks() call, healthCheckManager interface
  • hypershift-operator/main_test.go: new test for setupHealthChecks
  • cmd/install/assets/hypershift_operator.go: probes from /metrics:9000 to /healthz:8081 and /readyz:8081, add health port constant
  • cmd/install/assets/hypershift_operator_test.go: probe assertions

Conflict resolution: hypershift-operator/main.go had merge conflicts because release-4.22 uses a monolithic run() function while main/release-5.0 refactored it into sub-functions. Health probe additions placed into existing structure.

Refs OCPBUGS-115188


Commit 2: feat(control-plane-pki-operator): add TLS security profile configuration

Cherry-pick of 18352b86f1f (original PR #8768, CNTRLPLANE-3624 — clean, no conflicts).

Root cause: The e2e-v2-aws test suite (from main's hypershift-tests:latest) includes a blocking TLS test that expects a control-plane-pki-operator-config ConfigMap in the control plane namespace. This ConfigMap is created by CPOv2's controller-config.yaml asset with a manifest adapter in the pki-operator component — a feature added on main in June 2026 and present on release-5.0, but never backported to release-4.22. Without it, the test fails deterministically (0/12 pass rate, lifecycle: blocking).

Sippy data confirms: test [Feature:ControlPlaneTLS] ... control-plane-pki-operator-config ConfigMap with TLS configuration — 0% pass rate, 12 consecutive failures on release-4.22.

Changes:

  • v2/assets/control-plane-pki-operator/controller-config.yaml: new ConfigMap template
  • v2/assets/control-plane-pki-operator/deployment.yaml: mount config + --config and --terminate-on-files flags
  • v2/pkioperator/component.go: register WithManifestAdapter for controller-config.yaml
  • v2/pkioperator/configmap.go: adapter function deriving TLS settings from HCP spec
  • v2/pkioperator/configmap_test.go: tests for all TLS profile variants

The pki-operator binary on 4.22 already supports --config via the controllercmd framework from library-go.


Commit 3: chore: regenerate test data fixtures

Cherry-pick of 493721270c9 — regenerated testdata fixtures for TestControlPlaneComponents to include the new control-plane-pki-operator-config ConfigMap and updated deployment spec from commit 2.


Why all three in one PR

The e2e-v2-aws presubmit has two independent failure modes:

  • Pattern A (deterministic): TLS ConfigMap test fails → fixed by commits 2+3
  • Pattern B (most runs): upgrade step fails due to health probe mismatch → fixed by commit 1

Without commit 1, the upgrade step crashes and cascades into subsequent test failures.
Without commits 2+3, the TLS test fails even when tests execute successfully.

Neither fix alone makes e2e-v2-aws pass, so merging them separately creates a deadlock where neither PR can satisfy CI.

References

Test plan

  • go build ./hypershift-operator/... passes
  • go test ./hypershift-operator/ -run TestSetupHealthChecks passes
  • go test ./cmd/install/assets/ -run TestHyperShiftOperatorDeployment passes
  • go test ./control-plane-operator/controllers/hostedcontrolplane/v2/pkioperator/... passes (all 6 TLS profile tests)
  • TestControlPlaneComponents passes (fixtures regenerated)
  • e2e-aws-upgrade-hypershift-operator CI job passes
  • e2e-v2-aws CI job passes

🤖 Generated with Claude Code

Cherry-pick of 4084a52 from main/release-5.0 to release-4.22.
Conflict in hypershift-operator/main.go resolved: release-4.22 uses a
monolithic run() function whereas main/release-5.0 refactored it into
sub-functions (createManager, validateStartOptions, etc.). The health
probe additions were placed into the existing monolithic structure.

Root cause of CI failure (e2e-aws-upgrade-hypershift-operator 0/10):
The upgrade test binary (hypershift-tests:latest) is built from main,
not from the release branch. Commit 4084a52 on main changed the HO
deployment probes from /metrics:9000 to /healthz:8081 and /readyz:8081.
When the test upgrades the HO, it deploys with 8081 probes, but the
release-4.22 operator binary lacked HealthProbeBindAddress — no health
server on 8081 — readiness probe fails — rollout stuck at 1/2 replicas
— timeout after 608s.

Changes (identical to the original commit, adapted to 4.22 structure):
- hypershift-operator/main.go: add HealthProbeBindAddress to manager
  options, add setupHealthChecks() call after manager creation, add
  healthCheckManager interface and setupHealthChecks function
- hypershift-operator/main_test.go: new test for setupHealthChecks
- cmd/install/assets/hypershift_operator.go: change probes from
  /metrics:9000 to /healthz:8081 and /readyz:8081, add health port,
  add HypershiftOperatorHealthProbePort constant
- cmd/install/assets/hypershift_operator_test.go: add probe assertions

Original: openshift#9387 (OCPBUGS-111601, OCPBUGS-115188)
Backport to 5.0: openshift#9534
Refs OCPBUGS-115188

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Juan Manuel Parrilla Madrid <jparrill@redhat.com>
@openshift-ci

openshift-ci Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

Pipeline controller notification
This repo is configured to use the pipeline controller. Second-stage tests will be triggered either automatically or after lgtm label is added, depending on the repository configuration. The pipeline controller will automatically detect which contexts are required and will utilize /test Prow commands to trigger the second stage.

For optional jobs, comment /test ? to see a list of all defined jobs. To trigger manually all jobs from second stage use /pipeline required command.

This repository is configured in: LGTM mode

@coderabbitai

coderabbitai Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 410a5ca7-c61e-4031-a88e-824f0b90581e

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@openshift-ci
openshift-ci Bot requested review from bryan-cox and csrwng September 14, 2026 11:41
@openshift-ci openshift-ci Bot added area/cli Indicates the PR includes changes for CLI area/hypershift-operator Indicates the PR includes changes for the hypershift operator and API - outside an OCP release approved Indicates a PR has been approved by an approver from all required OWNERS files. and removed do-not-merge/needs-area labels Sep 14, 2026
@jparrill

Copy link
Copy Markdown
Contributor Author

/jira cherrypick OCPBUGS-115188

@openshift-ci-robot

Copy link
Copy Markdown

@jparrill: Detected clone of Jira Issue OCPBUGS-115188 with correct target version. Will retitle the PR to link to the clone.
/retitle OCPBUGS-120704: [release-4.22] fix(operator): use dedicated health probes

Details

In response to this:

/jira cherrypick OCPBUGS-115188

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-ci openshift-ci Bot changed the title [release-4.22] fix(operator): use dedicated health probes OCPBUGS-120704: [release-4.22] fix(operator): use dedicated health probes Sep 14, 2026
@openshift-ci-robot openshift-ci-robot added jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. jira/invalid-bug Indicates that a referenced Jira bug is invalid for the branch this PR is targeting. labels Sep 14, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@jparrill: This pull request references Jira Issue OCPBUGS-120704, which is invalid:

  • expected dependent Jira Issue OCPBUGS-115188 to be in one of the following states: MODIFIED, ON_QA, VERIFIED, RELEASE PENDING, CLOSED (ERRATA), CLOSED (CURRENT RELEASE), CLOSED (DONE), CLOSED (DONE-ERRATA), but it is POST instead

Comment /jira refresh to re-evaluate validity if changes to the Jira bug are made, or edit the title of this pull request to link to a different bug.

The bug has been updated to refer to the pull request using the external bug tracker.

Details

In response to this:

Summary

Cherry-pick of 4084a520609 to release-4.22, fixing e2e-aws-upgrade-hypershift-operator (0/10 BROKEN).

Root cause

The upgrade test binary (hypershift-tests:latest) is built from main, not from the release branch being tested. The CI config confirms this:

hypershift-tests:
   name: hypershift-tests
   namespace: hypershift
   tag: latest    # ← built from main

Commit 4084a520609 on main changed the HO deployment probes from /metrics:9000 to /healthz:8081 and /readyz:8081. When the upgrade test deploys the new HO, it uses main's install code which sets port 8081 probes. However, the release-4.22 operator binary lacked HealthProbeBindAddress — no health server listens on port 8081 — the readiness probe fails — rollout gets stuck at 1/2 replicas — the test times out after 608s.

Failure chain (empirically verified)

  1. CI runs e2e-aws-upgrade-hypershift-operator on release-4.22 PR
  2. Initial HO install uses 4.22's hcp CLI (INSTALL_FROM_LATEST=true) → probes at /metrics:9000 → works
  3. Upgrade step uses hypershift-tests:latest (from main) → deploys HO with /healthz:8081 and /readyz:8081 probes
  4. Release-4.22 operator binary has no HealthProbeBindAddress → port 8081 never opens
  5. Readiness probe fails → new replica never becomes ready → rollout stuck at 1/2
  6. Test timeout at 608s → FAIL

Changes

Identical to the original commit, adapted to release-4.22's monolithic run() structure (main/5.0 have refactored sub-functions):

  • hypershift-operator/main.go: Add HealthProbeBindAddress to manager options, add setupHealthChecks() call after manager creation, add healthCheckManager interface and setupHealthChecks function
  • hypershift-operator/main_test.go: New test for setupHealthChecks
  • cmd/install/assets/hypershift_operator.go: Change probes from /metrics:9000 to /healthz:8081 and /readyz:8081, add health port, add HypershiftOperatorHealthProbePort constant
  • cmd/install/assets/hypershift_operator_test.go: Add probe assertions

Conflict resolution

hypershift-operator/main.go had merge conflicts because release-4.22 uses a monolithic run() function, while main/release-5.0 refactored it into sub-functions (createManager, validateStartOptions, resolveOperatorImage, etc.). The health probe additions were placed into the existing monolithic structure:

  • HealthProbeBindAddress added to the inline ctrl.NewManager call
  • setupHealthChecks() call placed after manager creation inside run()
  • healthCheckManager interface and setupHealthChecks function added at end of file

References

Test plan

  • go build ./hypershift-operator/... passes
  • go test ./hypershift-operator/ -run TestSetupHealthChecks passes
  • go test ./cmd/install/assets/ -run TestHyperShiftOperatorDeployment passes
  • e2e-aws-upgrade-hypershift-operator CI job passes on this PR

🤖 Generated with Claude Code

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@jparrill

Copy link
Copy Markdown
Contributor Author

/jira refresh

@openshift-ci-robot openshift-ci-robot added jira/valid-bug Indicates that a referenced Jira bug is valid for the branch this PR is targeting. and removed jira/invalid-bug Indicates that a referenced Jira bug is invalid for the branch this PR is targeting. labels Sep 14, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@jparrill: This pull request references Jira Issue OCPBUGS-120704, which is valid. The bug has been moved to the POST state.

7 validation(s) were run on this bug
  • bug is open, matching expected state (open)
  • bug target version (4.22.0) matches configured target version for branch (4.22.0)
  • bug is in the state New, which is one of the valid states (NEW, ASSIGNED, POST)
  • release note type set to "Release Note Not Required"
  • dependent bug Jira Issue OCPBUGS-115188 is in the state MODIFIED, which is one of the valid states (MODIFIED, ON_QA, VERIFIED, RELEASE PENDING, CLOSED (ERRATA), CLOSED (CURRENT RELEASE), CLOSED (DONE), CLOSED (DONE-ERRATA))
  • dependent Jira Issue OCPBUGS-115188 targets the "5.0.0" version, which is one of the valid target versions: 5.0.0
  • bug has dependents
Details

In response to this:

/jira refresh

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@codecov

codecov Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 64.06250% with 23 lines in your changes missing coverage. Please review.
✅ Project coverage is 36.77%. Comparing base (7273000) to head (998cb0e).
⚠️ Report is 8 commits behind head on release-4.22.

Files with missing lines Patch % Lines
hypershift-operator/main.go 44.44% 8 Missing and 2 partials ⚠️
...ers/hostedcontrolplane/v2/pkioperator/configmap.go 72.72% 6 Missing and 3 partials ⚠️
...ers/hostedcontrolplane/v2/pkioperator/component.go 0.00% 4 Missing ⚠️
Additional details and impacted files
@@               Coverage Diff                @@
##           release-4.22    #9600      +/-   ##
================================================
+ Coverage         36.75%   36.77%   +0.01%     
================================================
  Files               779      780       +1     
  Lines             95579    95639      +60     
================================================
+ Hits              35130    35167      +37     
- Misses            57605    57623      +18     
- Partials           2844     2849       +5     
Files with missing lines Coverage Δ
cmd/install/assets/hypershift_operator.go 43.74% <100.00%> (+0.14%) ⬆️
...ers/hostedcontrolplane/v2/pkioperator/component.go 0.00% <0.00%> (ø)
...ers/hostedcontrolplane/v2/pkioperator/configmap.go 72.72% <72.72%> (ø)
hypershift-operator/main.go 1.08% <44.44%> (+1.08%) ⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@openshift-ci-robot

Copy link
Copy Markdown

@jparrill: This pull request references Jira Issue OCPBUGS-120704, which is valid.

7 validation(s) were run on this bug
  • bug is open, matching expected state (open)
  • bug target version (4.22.0) matches configured target version for branch (4.22.0)
  • bug is in the state POST, which is one of the valid states (NEW, ASSIGNED, POST)
  • release note type set to "Release Note Not Required"
  • dependent bug Jira Issue OCPBUGS-115188 is in the state MODIFIED, which is one of the valid states (MODIFIED, ON_QA, VERIFIED, RELEASE PENDING, CLOSED (ERRATA), CLOSED (CURRENT RELEASE), CLOSED (DONE), CLOSED (DONE-ERRATA))
  • dependent Jira Issue OCPBUGS-115188 targets the "5.0.0" version, which is one of the valid target versions: 5.0.0
  • bug has dependents
Details

In response to this:

Summary

Cherry-pick of 4084a520609 to release-4.22, fixing e2e-aws-upgrade-hypershift-operator (0/10 BROKEN).

Root cause

The upgrade test binary (hypershift-tests:latest) is built from main, not from the release branch being tested. The CI config confirms this:

hypershift-tests:
   name: hypershift-tests
   namespace: hypershift
   tag: latest    # ← built from main

Commit 4084a520609 on main changed the HO deployment probes from /metrics:9000 to /healthz:8081 and /readyz:8081. When the upgrade test deploys the new HO, it uses main's install code which sets port 8081 probes. However, the release-4.22 operator binary lacked HealthProbeBindAddress — no health server listens on port 8081 — the readiness probe fails — rollout gets stuck at 1/2 replicas — the test times out after 608s.

Fixes

Failure chain (empirically verified)

  1. CI runs e2e-aws-upgrade-hypershift-operator on release-4.22 PR
  2. Initial HO install uses 4.22's hcp CLI (INSTALL_FROM_LATEST=true) → probes at /metrics:9000 → works
  3. Upgrade step uses hypershift-tests:latest (from main) → deploys HO with /healthz:8081 and /readyz:8081 probes
  4. Release-4.22 operator binary has no HealthProbeBindAddress → port 8081 never opens
  5. Readiness probe fails → new replica never becomes ready → rollout stuck at 1/2
  6. Test timeout at 608s → FAIL

Changes

Identical to the original commit, adapted to release-4.22's monolithic run() structure (main/5.0 have refactored sub-functions):

  • hypershift-operator/main.go: Add HealthProbeBindAddress to manager options, add setupHealthChecks() call after manager creation, add healthCheckManager interface and setupHealthChecks function
  • hypershift-operator/main_test.go: New test for setupHealthChecks
  • cmd/install/assets/hypershift_operator.go: Change probes from /metrics:9000 to /healthz:8081 and /readyz:8081, add health port, add HypershiftOperatorHealthProbePort constant
  • cmd/install/assets/hypershift_operator_test.go: Add probe assertions

Conflict resolution

hypershift-operator/main.go had merge conflicts because release-4.22 uses a monolithic run() function, while main/release-5.0 refactored it into sub-functions (createManager, validateStartOptions, resolveOperatorImage, etc.). The health probe additions were placed into the existing monolithic structure:

  • HealthProbeBindAddress added to the inline ctrl.NewManager call
  • setupHealthChecks() call placed after manager creation inside run()
  • healthCheckManager interface and setupHealthChecks function added at end of file

References

Test plan

  • go build ./hypershift-operator/... passes
  • go test ./hypershift-operator/ -run TestSetupHealthChecks passes
  • go test ./cmd/install/assets/ -run TestHyperShiftOperatorDeployment passes
  • e2e-aws-upgrade-hypershift-operator CI job passes on this PR

🤖 Generated with Claude Code

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-ci openshift-ci Bot added the lgtm Indicates that a PR is ready to be merged. label Sep 14, 2026
@openshift-ci

openshift-ci Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

Scheduling tests matching the pipeline_run_if_changed or not excluded by pipeline_skip_if_only_changed parameters:
/test e2e-aks-4-21
/test e2e-aws-4-21
/test e2e-aks
/test e2e-aws
/test e2e-aws-upgrade-hypershift-operator
/test e2e-azure-aks-external-oidc
/test e2e-azure-self-managed
/test e2e-kubevirt-aws-ovn-reduced
/test e2e-v2-aws

@openshift-ci

openshift-ci Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: bryan-cox, jparrill

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@csrwng csrwng added the backport-risk-assessed Indicates a PR to a release branch has been evaluated and considered safe to accept. label Sep 14, 2026
configure the control-plane-pki-operator to use the tls security profile
settings from the hostedcontrolplane resource. this ensures the
operator's metrics endpoint uses ciphers and minimum tls version that
match the cluster's security requirements.

implementation:
- add configmap adapter to generate genericcontrollerconfig with tls
  settings derived from hcp.spec.configuration.tlssecurityprofile
- mount the config and uses --config flag on the operator deployment
- reuse existing config.ciphersuites() and config.mintlsversion()
  helper functions for consistency with other control plane components

this implementation is very similar to the registry operator
implementation.
@openshift-ci openshift-ci Bot removed the lgtm Indicates that a PR is ready to be merged. label Sep 14, 2026
@vsolanki12

Copy link
Copy Markdown
Contributor

/retest

@red-hat-konflux

Copy link
Copy Markdown
Contributor

All PipelineRuns for this commit have already succeeded. Use /retest <pipeline-name> to re-run a specific pipeline or /test to re-run all pipelines.

@jparrill

Copy link
Copy Markdown
Contributor Author

/retest-required

@jparrill

Copy link
Copy Markdown
Contributor Author

/test e2e-v2-aws

@sdminonne

Copy link
Copy Markdown
Contributor

/test e2e-aws-upgrade-hypershift-operator

@bryan-cox

Copy link
Copy Markdown
Member

/test e2e-v2-aws

Looks like a test flake with one failing test

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

/retest-required

Remaining retests: 0 against base HEAD 211f9b0 and 1 for PR HEAD 998cb0e in total

@vsolanki12

Copy link
Copy Markdown
Contributor

/test e2e-aws-upgrade-hypershift-operator
/test e2e-v2-aws

@jparrill

jparrill commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

/retest-required

Remaining retests: 0 against base HEAD 5f96792 and 0 for PR HEAD 998cb0e in total

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

/hold

Revision 998cb0e was retested 3 times: holding

@openshift-ci openshift-ci Bot added the do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. label Sep 17, 2026
@vsolanki12

Copy link
Copy Markdown
Contributor

/unhold
/test e2e-v2-aws

@openshift-ci openshift-ci Bot removed the do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. label Sep 18, 2026
@vsolanki12

Copy link
Copy Markdown
Contributor

/test e2e-aws-upgrade-hypershift-operator

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

/retest-required

Remaining retests: 0 against base HEAD 5f96792 and 2 for PR HEAD 998cb0e in total

@jparrill

Copy link
Copy Markdown
Contributor Author

/test e2e-aws-upgrade-hypershift-operator

1 similar comment
@jparrill

Copy link
Copy Markdown
Contributor Author

/test e2e-aws-upgrade-hypershift-operator

@jparrill

Copy link
Copy Markdown
Contributor Author

/override ci/prow/e2e-aws-upgrade-hypershift-operator

Too much time with this blocking issue, we cannot wait more for a fix for the above issue. It should be solved via #9736 but still on testing...

@openshift-ci

openshift-ci Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

@jparrill: Overrode contexts on behalf of jparrill: ci/prow/e2e-aws-upgrade-hypershift-operator

Details

In response to this:

/override ci/prow/e2e-aws-upgrade-hypershift-operator

Too much time with this blocking issue, we cannot wait more for a fix for the above issue. It should be solved via #9736 but still on testing...

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@openshift-ci

openshift-ci Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

@jparrill: all tests passed!

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@openshift-merge-bot
openshift-merge-bot Bot merged commit d516aea into openshift:release-4.22 Sep 23, 2026
28 checks passed
@openshift-ci-robot

Copy link
Copy Markdown

@jparrill: Jira Issue Verification Checks: Jira Issue OCPBUGS-120704
✔️ This pull request was pre-merge verified.
✔️ All associated pull requests have merged.
✔️ All associated, merged pull requests were pre-merge verified.

Jira Issue OCPBUGS-120704 has been moved to the MODIFIED state and will move to the VERIFIED state when the change is available in an accepted nightly payload. 🕓

Details

In response to this:

Summary

Three cherry-picks to release-4.22 that fix both e2e-aws-upgrade-hypershift-operator and e2e-v2-aws presubmits (both 0/10 BROKEN, blocking the merge pipeline).

Both functional fixes are needed in the same PR to avoid a chicken-and-egg deadlock: neither presubmit can pass without both changes present, so merging them separately is impossible.


Commit 1: fix(operator): use dedicated health probes

Cherry-pick of 4084a520609 (original PR #9387, backported to 5.0 via #9534).

Root cause: The upgrade test binary (hypershift-tests:latest) is built from main, not from the release branch. Main's install code (cmd/install/assets/hypershift_operator.go) changed HO deployment probes from /metrics:9000 to /healthz:8081 and /readyz:8081. When the upgrade test deploys the new HO, it sets port 8081 probes, but the release-4.22 operator binary lacked HealthProbeBindAddress — no health server on 8081 — readiness probe fails — rollout stuck at 1/2 replicas — timeout after 608s.

Failure chain (empirically verified):

  1. CI runs e2e-aws-upgrade-hypershift-operator on release-4.22 PR
  2. Initial HO install uses 4.22's hcp CLI (INSTALL_FROM_LATEST=true) → probes at /metrics:9000 → works
  3. Upgrade step uses hypershift-tests:latest (from main) → deploys HO with /healthz:8081 and /readyz:8081 probes
  4. Release-4.22 operator binary has no HealthProbeBindAddress → port 8081 never opens
  5. Readiness probe fails → new replica never becomes ready → rollout stuck at 1/2
  6. Test timeout at 608s → FAIL

Changes (adapted to 4.22's monolithic run() structure):

  • hypershift-operator/main.go: add HealthProbeBindAddress, setupHealthChecks() call, healthCheckManager interface
  • hypershift-operator/main_test.go: new test for setupHealthChecks
  • cmd/install/assets/hypershift_operator.go: probes from /metrics:9000 to /healthz:8081 and /readyz:8081, add health port constant
  • cmd/install/assets/hypershift_operator_test.go: probe assertions

Conflict resolution: hypershift-operator/main.go had merge conflicts because release-4.22 uses a monolithic run() function while main/release-5.0 refactored it into sub-functions. Health probe additions placed into existing structure.

Refs OCPBUGS-115188


Commit 2: feat(control-plane-pki-operator): add TLS security profile configuration

Cherry-pick of 18352b86f1f (original PR #8768, CNTRLPLANE-3624 — clean, no conflicts).

Root cause: The e2e-v2-aws test suite (from main's hypershift-tests:latest) includes a blocking TLS test that expects a control-plane-pki-operator-config ConfigMap in the control plane namespace. This ConfigMap is created by CPOv2's controller-config.yaml asset with a manifest adapter in the pki-operator component — a feature added on main in June 2026 and present on release-5.0, but never backported to release-4.22. Without it, the test fails deterministically (0/12 pass rate, lifecycle: blocking).

Sippy data confirms: test [Feature:ControlPlaneTLS] ... control-plane-pki-operator-config ConfigMap with TLS configuration — 0% pass rate, 12 consecutive failures on release-4.22.

Changes:

  • v2/assets/control-plane-pki-operator/controller-config.yaml: new ConfigMap template
  • v2/assets/control-plane-pki-operator/deployment.yaml: mount config + --config and --terminate-on-files flags
  • v2/pkioperator/component.go: register WithManifestAdapter for controller-config.yaml
  • v2/pkioperator/configmap.go: adapter function deriving TLS settings from HCP spec
  • v2/pkioperator/configmap_test.go: tests for all TLS profile variants

The pki-operator binary on 4.22 already supports --config via the controllercmd framework from library-go.


Commit 3: chore: regenerate test data fixtures

Cherry-pick of 493721270c9 — regenerated testdata fixtures for TestControlPlaneComponents to include the new control-plane-pki-operator-config ConfigMap and updated deployment spec from commit 2.


Why all three in one PR

The e2e-v2-aws presubmit has two independent failure modes:

  • Pattern A (deterministic): TLS ConfigMap test fails → fixed by commits 2+3
  • Pattern B (most runs): upgrade step fails due to health probe mismatch → fixed by commit 1

Without commit 1, the upgrade step crashes and cascades into subsequent test failures.
Without commits 2+3, the TLS test fails even when tests execute successfully.

Neither fix alone makes e2e-v2-aws pass, so merging them separately creates a deadlock where neither PR can satisfy CI.

References

Test plan

  • go build ./hypershift-operator/... passes
  • go test ./hypershift-operator/ -run TestSetupHealthChecks passes
  • go test ./cmd/install/assets/ -run TestHyperShiftOperatorDeployment passes
  • go test ./control-plane-operator/controllers/hostedcontrolplane/v2/pkioperator/... passes (all 6 TLS profile tests)
  • TestControlPlaneComponents passes (fixtures regenerated)
  • e2e-aws-upgrade-hypershift-operator CI job passes
  • e2e-v2-aws CI job passes

🤖 Generated with Claude Code

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@sdminonne

Copy link
Copy Markdown
Contributor

/cherrypick release-4.21

@openshift-cherrypick-robot

Copy link
Copy Markdown

@sdminonne: #9600 failed to apply on top of branch "release-4.21":

Applying: fix(operator): use dedicated health probes
Using index info to reconstruct a base tree...
M	cmd/install/assets/hypershift_operator.go
M	cmd/install/assets/hypershift_operator_test.go
M	hypershift-operator/main.go
Falling back to patching base and 3-way merge...
Auto-merging cmd/install/assets/hypershift_operator.go
Auto-merging cmd/install/assets/hypershift_operator_test.go
Auto-merging hypershift-operator/main.go
CONFLICT (content): Merge conflict in hypershift-operator/main.go
error: Failed to merge in the changes.
hint: Use 'git am --show-current-patch=diff' to see the failed patch
hint: When you have resolved this problem, run "git am --continue".
hint: If you prefer to skip this patch, run "git am --skip" instead.
hint: To restore the original branch and stop patching, run "git am --abort".
hint: Disable this message with "git config set advice.mergeConflict false"
Patch failed at 0001 fix(operator): use dedicated health probes

Details

In response to this:

/cherrypick release-4.21

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. area/cli Indicates the PR includes changes for CLI area/control-plane-operator Indicates the PR includes changes for the control plane operator - in an OCP release area/hypershift-operator Indicates the PR includes changes for the hypershift operator and API - outside an OCP release backport-risk-assessed Indicates a PR to a release branch has been evaluated and considered safe to accept. jira/valid-bug Indicates that a referenced Jira bug is valid for the branch this PR is targeting. jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. lgtm Indicates that a PR is ready to be merged. verified Signifies that the PR passed pre-merge verification criteria

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants