Repository navigation
OCPBUGS-120704: [release-4.22] fix(operator): health probes + pki-operator TLS config - #9600
Conversation
Cherry-pick of 4084a52 from main/release-5.0 to release-4.22. Conflict in hypershift-operator/main.go resolved: release-4.22 uses a monolithic run() function whereas main/release-5.0 refactored it into sub-functions (createManager, validateStartOptions, etc.). The health probe additions were placed into the existing monolithic structure. Root cause of CI failure (e2e-aws-upgrade-hypershift-operator 0/10): The upgrade test binary (hypershift-tests:latest) is built from main, not from the release branch. Commit 4084a52 on main changed the HO deployment probes from /metrics:9000 to /healthz:8081 and /readyz:8081. When the test upgrades the HO, it deploys with 8081 probes, but the release-4.22 operator binary lacked HealthProbeBindAddress — no health server on 8081 — readiness probe fails — rollout stuck at 1/2 replicas — timeout after 608s. Changes (identical to the original commit, adapted to 4.22 structure): - hypershift-operator/main.go: add HealthProbeBindAddress to manager options, add setupHealthChecks() call after manager creation, add healthCheckManager interface and setupHealthChecks function - hypershift-operator/main_test.go: new test for setupHealthChecks - cmd/install/assets/hypershift_operator.go: change probes from /metrics:9000 to /healthz:8081 and /readyz:8081, add health port, add HypershiftOperatorHealthProbePort constant - cmd/install/assets/hypershift_operator_test.go: add probe assertions Original: openshift#9387 (OCPBUGS-111601, OCPBUGS-115188) Backport to 5.0: openshift#9534 Refs OCPBUGS-115188 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Juan Manuel Parrilla Madrid <jparrill@redhat.com>
|
Pipeline controller notification For optional jobs, comment This repository is configured in: LGTM mode |
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository YAML (base), Central YAML (inherited) Review profile: CHILL Plan: Enterprise Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
|
/jira cherrypick OCPBUGS-115188 |
|
@jparrill: Detected clone of Jira Issue OCPBUGS-115188 with correct target version. Will retitle the PR to link to the clone. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
@jparrill: This pull request references Jira Issue OCPBUGS-120704, which is invalid:
Comment The bug has been updated to refer to the pull request using the external bug tracker. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
/jira refresh |
|
@jparrill: This pull request references Jira Issue OCPBUGS-120704, which is valid. The bug has been moved to the POST state. 7 validation(s) were run on this bug
DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## release-4.22 #9600 +/- ##
================================================
+ Coverage 36.75% 36.77% +0.01%
================================================
Files 779 780 +1
Lines 95579 95639 +60
================================================
+ Hits 35130 35167 +37
- Misses 57605 57623 +18
- Partials 2844 2849 +5
🚀 New features to boost your workflow:
|
|
@jparrill: This pull request references Jira Issue OCPBUGS-120704, which is valid. 7 validation(s) were run on this bug
DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
Scheduling tests matching the |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: bryan-cox, jparrill The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
configure the control-plane-pki-operator to use the tls security profile settings from the hostedcontrolplane resource. this ensures the operator's metrics endpoint uses ciphers and minimum tls version that match the cluster's security requirements. implementation: - add configmap adapter to generate genericcontrollerconfig with tls settings derived from hcp.spec.configuration.tlssecurityprofile - mount the config and uses --config flag on the operator deployment - reuse existing config.ciphersuites() and config.mintlsversion() helper functions for consistency with other control plane components this implementation is very similar to the registry operator implementation.
|
/retest |
|
All PipelineRuns for this commit have already succeeded. Use |
|
/retest-required |
|
/test e2e-v2-aws |
|
/test e2e-aws-upgrade-hypershift-operator |
|
/test e2e-v2-aws Looks like a test flake with one failing test |
|
/test e2e-aws-upgrade-hypershift-operator |
|
We need this for E2E to pass |
|
/hold Revision 998cb0e was retested 3 times: holding |
|
/unhold |
|
/test e2e-aws-upgrade-hypershift-operator |
|
/test e2e-aws-upgrade-hypershift-operator |
1 similar comment
|
/test e2e-aws-upgrade-hypershift-operator |
|
/override ci/prow/e2e-aws-upgrade-hypershift-operator Too much time with this blocking issue, we cannot wait more for a fix for the above issue. It should be solved via #9736 but still on testing... |
|
@jparrill: Overrode contexts on behalf of jparrill: ci/prow/e2e-aws-upgrade-hypershift-operator DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
@jparrill: all tests passed! Full PR test history. Your PR dashboard. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here. |
d516aea
into
openshift:release-4.22
|
@jparrill: Jira Issue Verification Checks: Jira Issue OCPBUGS-120704 Jira Issue OCPBUGS-120704 has been moved to the MODIFIED state and will move to the VERIFIED state when the change is available in an accepted nightly payload. 🕓 DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
/cherrypick release-4.21 |
|
@sdminonne: #9600 failed to apply on top of branch "release-4.21": DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
Summary
Three cherry-picks to release-4.22 that fix both
e2e-aws-upgrade-hypershift-operatorande2e-v2-awspresubmits (both 0/10 BROKEN, blocking the merge pipeline).Both functional fixes are needed in the same PR to avoid a chicken-and-egg deadlock: neither presubmit can pass without both changes present, so merging them separately is impossible.
Commit 1: fix(operator): use dedicated health probes
Cherry-pick of
4084a520609(original PR #9387, backported to 5.0 via #9534).Root cause: The upgrade test binary (
hypershift-tests:latest) is built from main, not from the release branch. Main's install code (cmd/install/assets/hypershift_operator.go) changed HO deployment probes from/metrics:9000to/healthz:8081and/readyz:8081. When the upgrade test deploys the new HO, it sets port 8081 probes, but the release-4.22 operator binary lackedHealthProbeBindAddress— no health server on 8081 — readiness probe fails — rollout stuck at 1/2 replicas — timeout after 608s.Failure chain (empirically verified):
e2e-aws-upgrade-hypershift-operatoron release-4.22 PRhcpCLI (INSTALL_FROM_LATEST=true) → probes at/metrics:9000→ workshypershift-tests:latest(from main) → deploys HO with/healthz:8081and/readyz:8081probesHealthProbeBindAddress→ port 8081 never opensChanges (adapted to 4.22's monolithic
run()structure):hypershift-operator/main.go: addHealthProbeBindAddress,setupHealthChecks()call,healthCheckManagerinterfacehypershift-operator/main_test.go: new test forsetupHealthCheckscmd/install/assets/hypershift_operator.go: probes from/metrics:9000to/healthz:8081and/readyz:8081, add health port constantcmd/install/assets/hypershift_operator_test.go: probe assertionsConflict resolution:
hypershift-operator/main.gohad merge conflicts because release-4.22 uses a monolithicrun()function while main/release-5.0 refactored it into sub-functions. Health probe additions placed into existing structure.Refs OCPBUGS-115188
Commit 2: feat(control-plane-pki-operator): add TLS security profile configuration
Cherry-pick of
18352b86f1f(original PR #8768, CNTRLPLANE-3624 — clean, no conflicts).Root cause: The
e2e-v2-awstest suite (from main'shypershift-tests:latest) includes a blocking TLS test that expects acontrol-plane-pki-operator-configConfigMap in the control plane namespace. This ConfigMap is created by CPOv2'scontroller-config.yamlasset with a manifest adapter in the pki-operator component — a feature added on main in June 2026 and present on release-5.0, but never backported to release-4.22. Without it, the test fails deterministically (0/12 pass rate, lifecycle: blocking).Sippy data confirms: test
[Feature:ControlPlaneTLS] ... control-plane-pki-operator-config ConfigMap with TLS configuration— 0% pass rate, 12 consecutive failures on release-4.22.Changes:
v2/assets/control-plane-pki-operator/controller-config.yaml: new ConfigMap templatev2/assets/control-plane-pki-operator/deployment.yaml: mount config +--configand--terminate-on-filesflagsv2/pkioperator/component.go: registerWithManifestAdapterforcontroller-config.yamlv2/pkioperator/configmap.go: adapter function deriving TLS settings from HCP specv2/pkioperator/configmap_test.go: tests for all TLS profile variantsThe pki-operator binary on 4.22 already supports
--configvia thecontrollercmdframework from library-go.Commit 3: chore: regenerate test data fixtures
Cherry-pick of
493721270c9— regenerated testdata fixtures forTestControlPlaneComponentsto include the newcontrol-plane-pki-operator-configConfigMap and updated deployment spec from commit 2.Why all three in one PR
The
e2e-v2-awspresubmit has two independent failure modes:Without commit 1, the upgrade step crashes and cascades into subsequent test failures.
Without commits 2+3, the TLS test fails even when tests execute successfully.
Neither fix alone makes
e2e-v2-awspass, so merging them separately creates a deadlock where neither PR can satisfy CI.References
493721270c9Test plan
go build ./hypershift-operator/...passesgo test ./hypershift-operator/ -run TestSetupHealthCheckspassesgo test ./cmd/install/assets/ -run TestHyperShiftOperatorDeploymentpassesgo test ./control-plane-operator/controllers/hostedcontrolplane/v2/pkioperator/...passes (all 6 TLS profile tests)TestControlPlaneComponentspasses (fixtures regenerated)e2e-aws-upgrade-hypershift-operatorCI job passese2e-v2-awsCI job passes🤖 Generated with Claude Code