Repository navigation
CNTRLPLANE-3434: hypershift-install: use mounted CI pull secret in extract_hcp_cli - #82061
Conversation
Follow-up fix for PR openshift#81877 which introduced extract_hcp_cli. The function used oc extract secret/pull-secret from the openshift-config namespace to authenticate with container registries. This fails on AKS clusters where openshift-config does not exist. Switch to the mounted CI pull secret at /etc/ci-pull-credentials/.dockerconfigjson, which is already used by the hypershift install command itself for Azure, GCP, and default cloud providers in the same script. Signed-off-by: Alessandro Rossi <alesross@redhat.com>
|
@Nirshal: This pull request references CNTRLPLANE-3434 which is a valid jira issue. Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the epic to target the "5.0.0" version, but no target version was set. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository YAML (base), Central YAML (inherited) Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
WalkthroughChangesHyperShift install updates
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 14 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (14 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.19-periodics-e2e-aks periodic-ci-openshift-hypershift-main-periodics-e2e-aws-multi periodic-ci-openshift-hypershift-main-e2e-aws-upgrade-hypershift-operator periodic-ci-openshift-hypershift-release-5.0-periodics-e2e-aks-upgrade-minor periodic-ci-openshift-hypershift-main-periodics-e2e-aws-install-from-latest |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/pj-rehearse pull-ci-openshift-hypershift-main-e2e-aks pull-ci-openshift-hypershift-main-e2e-aws pull-ci-openshift-hypershift-main-e2e-azure-v2-self-managed periodic-ci-openshift-hypershift-release-5.0-periodics-e2e-aks-multi-x-ax pull-ci-openshift-hypershift-main-e2e-aws-upgrade-hypershift-operator |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/pj-rehearse pull-ci-openshift-hypershift-main-e2e-aks |
|
@csrwng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
Trivial documentation update to trigger pj-rehearse detection. pj-rehearse does not detect isolated script-only changes without a corresponding ref YAML modification. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.19-periodics-e2e-aks periodic-ci-openshift-hypershift-main-periodics-e2e-aws-multi periodic-ci-openshift-hypershift-main-e2e-aws-upgrade-hypershift-operator periodic-ci-openshift-hypershift-release-5.0-periodics-e2e-aks-upgrade-minor periodic-ci-openshift-hypershift-main-periodics-e2e-aws-install-from-latest |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
@Nirshal: job(s): periodic-ci-openshift-hypershift-main-periodics-e2e-aws-multi, periodic-ci-openshift-hypershift-main-e2e-aws-upgrade-hypershift-operator, periodic-ci-openshift-hypershift-release-5.0-periodics-e2e-aks-upgrade-minor, periodic-ci-openshift-hypershift-main-periodics-e2e-aws-install-from-latest either don't exist or were not found to be affected, and cannot be rehearsed |
|
/pj-rehearse pull-ci-openshift-hypershift-main-e2e-aws pull-ci-openshift-hypershift-main-e2e-aks pull-ci-openshift-hypershift-main-e2e-aws-upgrade-hypershift-operator pull-ci-openshift-hypershift-main-e2e-azure-v2-self-managed periodic-ci-openshift-hypershift-release-5.0-periodics-e2e-aks-multi-x-ax |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/retest |
|
/pj-rehearse pull-ci-openshift-hypershift-main-e2e-aws pull-ci-openshift-hypershift-main-e2e-aks pull-ci-openshift-hypershift-main-e2e-azure-v2-self-managed periodic-ci-openshift-hypershift-release-5.0-periodics-e2e-aks-multi-x-ax |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/pj-rehearse pull-ci-openshift-hypershift-main-e2e-azure-v2-self-managed periodic-ci-openshift-hypershift-release-5.0-periodics-e2e-aks-multi-x-ax |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
Test Failure Analysis:
|
| Time (UTC) | Event |
|---|---|
| 08:46:05 | Pod created |
| 08:46:08 | Init container availability-prober completed |
| ~08:46:09 | 1st start → crash (API server 503) |
| 08:47:41 | 2nd start → crash at 08:48:03 (same error, exit code 1) |
| 08:48:13 | 3rd start → stable (running at dump time) |
The test asserts restartCount <= 0; it was 2. All other 25+ control plane workloads (kube-apiserver, etcd, kube-controller-manager, kube-scheduler, etc.) passed the same "no crashing pods" check with 0 restarts.
Proof This Is Unrelated to PR Changes
The hypershift-install step log (set -eux trace) shows the PR's new code paths were never executed in this job:
+ [[ -n '' ]] # MULTISTAGE_PARAM_OVERRIDE_OVERRIDE_HYPERSHIFT_OPERATOR_IMAGE — empty
+ [[ -n '' ]] # OVERRIDE_HYPERSHIFT_OPERATOR_IMAGE — empty
+ [[ false == true ]] # HO_MULTI — false
+ [[ false == true ]] # INSTALL_FROM_LATEST — falseAll three override flags evaluated to empty/false → extract_hcp_cli() was never called → HCP_CLI remained bin/hypershift (the default step container binary). The hypershift install command executed identically to what it would have been without this PR. The pull secret source (--pull-secret=/etc/ci-pull-credentials/.dockerconfigjson) was also unchanged.
Secondary Post-Step Failures
dump-management-cluster: management cluster API still unreachable after guest teardowndestroy-management-cluster:ImagePullBackOfffor CI image (manifest unknownin quay-proxy) — pod never started, timed out after 1h
Both are infrastructure issues unrelated to the test or the PR.
Analysis generated from Prow job artifacts
|
/pj-rehearse pull-ci-openshift-hypershift-main-e2e-azure-v2-self-managed |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/pj-rehearse pull-ci-openshift-hypershift-main-e2e-azure-v2-self-managed |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
Test Failure Analysis:
|
🔍 Test Failure Analysis — 2nd run (
|
| Phase | Step | Duration | Result |
|---|---|---|---|
| pre | ipi-install-rbac |
9s | ✅ |
| pre | create-management-cluster |
20m39s | ✅ |
| pre | hypershift-azure-setup-private-link |
57s | ✅ |
| pre | hypershift-install (modified by this PR) |
2m15s | ✅ |
| pre | hypershift-resolve-nodepool-releases |
35s | ✅ |
| pre | hypershift-azure-create-selfmanaged-guests |
17m28s | ✅ |
| test | tests |
1h15m6s | ✅ All tests passed |
| post | dump |
16m32s | ✅ |
| post | hypershift-debug |
10s | ✅ |
| post | hypershift-k8sgpt |
26s | ✅ |
| post | destroy-guests |
1h8m26s | ❌ Azure ResourceGroupDeletionTimeout |
| post | dump-management-cluster |
1h0m22s | ❌ ImagePullBackOff (manifest unknown) |
| post | destroy-management-cluster |
9m25s | ❌ Azure HTTP/2 stream cancelled |
Every pre and test phase step succeeded. Only 3 post-phase cleanup steps failed.
5. Root causes of the 3 post-phase failures
destroy-guests — Azure ARM API returned HTTP 409 ResourceGroupDeletionTimeout:
{
"code": "ResourceGroupDeletionTimeout",
"message": "Deletion of resource group 'autoscaling-56c02dfa17-nsg-autoscaling-56c02dfa1-xm7ht'
did not finish within the allowed time..."
}This is an Azure-side timeout deleting the NSG resource group for the autoscaling guest cluster.
dump-management-cluster — CI step image sha256:ce4bea79... doesn't exist in the registry:
manifest unknown in quay-proxy.ci.openshift.org/openshift/ci
manifest unknown in quay.io/openshift/ci (mirror)
The pod was pending for 1h with 260× Back-off pulling image events before being killed. This image is the step container for dump-management-cluster, resolved by ci-operator at runtime — it is not referenced or modified by this PR.
destroy-management-cluster — Azure HTTP/2 stream error during NSG resource group deletion:
stream error: stream ID 89; CANCEL; received from peer
6. PR changes are provably not involved
- The SHA
ce4bea79...(the failing image) does not appear anywhere in the PR diff:$ git diff main -- ci-operator/step-registry/hypershift/install/ | grep -c "ce4bea79" 0 - This PR only modifies
hypershift-install-commands.shandhypershift-install-ref.yaml. - The failing steps (
destroy-guests,dump-management-cluster,destroy-management-cluster) are completely separate step-registry components. - All three failures are caused by external infrastructure: Azure ARM API timeouts and CI image registry unavailability.
Conclusion
This is a pure infrastructure flake. The job's test phase completed with all tests passing. The failure is entirely in post-phase cleanup due to Azure resource group deletion timeouts and a missing CI step image in quay-proxy. None of these issues are related to the PR's changes to pull secret handling in extract_hcp_cli.
Recommendation: retry.
|
/pj-rehearse pull-ci-openshift-hypershift-main-e2e-azure-v2-self-managed |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/pj-rehearse ack |
|
@Nirshal: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: bryan-cox, Nirshal The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
a7c6ed7
into
openshift:main
…tract_hcp_cli (openshift#82061) * fix(hypershift-install): use mounted CI pull secret in extract_hcp_cli Follow-up fix for PR openshift#81877 which introduced extract_hcp_cli. The function used oc extract secret/pull-secret from the openshift-config namespace to authenticate with container registries. This fails on AKS clusters where openshift-config does not exist. Switch to the mounted CI pull secret at /etc/ci-pull-credentials/.dockerconfigjson, which is already used by the hypershift install command itself for Azure, GCP, and default cloud providers in the same script. Signed-off-by: Alessandro Rossi <alesross@redhat.com> * docs(hypershift-install): add HO release controller note to env doc Trivial documentation update to trigger pj-rehearse detection. pj-rehearse does not detect isolated script-only changes without a corresponding ref YAML modification. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Signed-off-by: Alessandro Rossi <alesross@redhat.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
…tract_hcp_cli (openshift#82061) * fix(hypershift-install): use mounted CI pull secret in extract_hcp_cli Follow-up fix for PR openshift#81877 which introduced extract_hcp_cli. The function used oc extract secret/pull-secret from the openshift-config namespace to authenticate with container registries. This fails on AKS clusters where openshift-config does not exist. Switch to the mounted CI pull secret at /etc/ci-pull-credentials/.dockerconfigjson, which is already used by the hypershift install command itself for Azure, GCP, and default cloud providers in the same script. Signed-off-by: Alessandro Rossi <alesross@redhat.com> * docs(hypershift-install): add HO release controller note to env doc Trivial documentation update to trigger pj-rehearse detection. pj-rehearse does not detect isolated script-only changes without a corresponding ref YAML modification. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Signed-off-by: Alessandro Rossi <alesross@redhat.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
…tract_hcp_cli (openshift#82061) * fix(hypershift-install): use mounted CI pull secret in extract_hcp_cli Follow-up fix for PR openshift#81877 which introduced extract_hcp_cli. The function used oc extract secret/pull-secret from the openshift-config namespace to authenticate with container registries. This fails on AKS clusters where openshift-config does not exist. Switch to the mounted CI pull secret at /etc/ci-pull-credentials/.dockerconfigjson, which is already used by the hypershift install command itself for Azure, GCP, and default cloud providers in the same script. Signed-off-by: Alessandro Rossi <alesross@redhat.com> * docs(hypershift-install): add HO release controller note to env doc Trivial documentation update to trigger pj-rehearse detection. pj-rehearse does not detect isolated script-only changes without a corresponding ref YAML modification. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Signed-off-by: Alessandro Rossi <alesross@redhat.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
What this PR does:
Switches
extract_hcp_cliinhypershift-install-commands.shfrom extracting the pull secret viaoc extract secret/pull-secret -n openshift-configto using the already-mounted CI pull secret at/etc/ci-pull-credentials/.dockerconfigjson.Why we need it:
Follow-up fix for #81877, which refactored the CLI extraction logic into a shared
extract_hcp_clifunction and addedOVERRIDE_HYPERSHIFT_OPERATOR_IMAGEsupport.The
openshift-confignamespace does not exist on AKS clusters, causingextract_hcp_clito fail when the HO release gate pipeline triggers AKS jobs with an image override.The mounted CI pull secret is already used by
hypershift install --pull-secretfor Azure, GCP, and default cloud providers in the same script. This change alignsextract_hcp_cliwith that existing pattern.Background:
The inline extraction logic was originally introduced for
HO_MULTI(bf81fdf, May 2024) andINSTALL_FROM_LATEST(3b34fa3, Nov 2024), both targeting OpenShift-only jobs whereopenshift-configalways exists. The newOVERRIDE_HYPERSHIFT_OPERATOR_IMAGEpath runs across all platforms including AKS, exposing the issue for the first time.Summary by CodeRabbit
hypershiftCLI using the mounted/etc/ci-pull-credentials/.dockerconfigjson, fixing compatibility with AKS clusters whereopenshift-configmay be unavailable.OVERRIDE_HYPERSHIFT_OPERATOR_IMAGEis used by the HO release controller.