Repository navigation
OCPSTRAT-3250: Konflux release gating pipeline for HyperShift Operator - #2016
Conversation
|
@bryan-cox: This pull request references OCPSTRAT-3250 which is a valid jira issue. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
e8dd6e1 to
68fbb5f
Compare
|
While we are implementing this effort for ARO HCP first, we are expecting to onboard ROSA HCP and GCP in the future. I wanted to make sure y'all were aware of this enhancement; please feel free to unsubscribe if you wish - @deads2k @joshbranham @cblecker |
5f6b081 to
a135a21
Compare
Introduces a nightly, platform-independent gating system that validates HyperShift Operator images against e2e test suites before promoting them to verified repositories. The pipeline operates alongside the existing Konflux auto-release, adding a parallel promotion path per managed service platform (ARO HCP pilot, ROSA HCP and GCP HCP future). OCPSTRAT-3250 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
dropped some more questions, lgtm |
- Add Design Rationale section capturing nightly cadence and alerting rationale from PR discussion threads (enxebre) - Clarify per-platform pipeline naming strategy with TBD for shared vs separate files when second platform onboards (enxebre) - Add Verified Repository to glossary as single shared quay.io repo with per-platform image tags (enxebre) - Document resource management: pipeline in .tekton/pipelines/, Konflux namespace resources in contrib/konflux/ (enxebre) - Fix CronJob/ITS interaction to follow Konflux periodic test pattern: ITS uses disabled context, CronJob labels snapshots instead of creating PipelineRuns directly (Nirshal) - Update RBAC, diagrams, and workflow steps to match new pattern Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
/approve |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: enxebre The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
JoelSpeed
left a comment
There was a problem hiding this comment.
One thing I'm hoping we solve as part of this EP is the ability to have confidence in the supported matrix of CPO to guest and CPO to management version skews. There is only a very light mention of cross version testing here, was that something you were considering in/out of scope?
|
|
||
| ## Summary | ||
|
|
||
| This enhancement introduces a nightly, platform-independent gating system that validates HyperShift Operator (HO) images against end-to-end test suites before promoting them to verified repositories. The pipeline operates alongside the existing Konflux auto-release mechanism, adding a parallel promotion path that only advances images which have passed e2e test suites agreed upon between the HyperShift and managed service (HCM) teams. Tests may vary by platform. |
There was a problem hiding this comment.
Promotion requires all tests pass across all platforms? Or are there separate promotion destinations such that we might see promotion succeed on ARO but not ROSA?
There was a problem hiding this comment.
There will be different promotion paths for each managed service since each one needs a different set of tests to pass. If ARO HCP tests fail but GCP and ROSA tests pass, they should still get a tagged HO for their managed services respectively.
There was a problem hiding this comment.
How do you see that increasing complexity for the HyperShift team? You'd end up with a partially green nightly rather than a full green or full red. It brings in complexity like having to understand "ok this PR went out to GCP and AWS but not Azure because XYZ"
I wonder if it's desirable from the HyperShift team perspective to just have a single signal for green vs red?
There was a problem hiding this comment.
I think the benefit of not blocking releases for issues that affect a single managed platform justifies having split reporting.
|
|
||
| #### Design Rationale | ||
|
|
||
| **Nightly cadence (24h):** Each pipeline run provisions real cloud infrastructure with platform-specific credentials (e.g., Azure for ARO HCP). A nightly cadence balances validation confidence with cloud infrastructure cost. Per-commit gating is cost-prohibitive and would slow the development feedback loop (see Alternatives). Per-platform cadence can differ — each platform can have its own CronJob schedule. |
There was a problem hiding this comment.
OpenShift CI and nightly builds happen every 6h, have you considered making this more frequent than once per day? Is there enough change in a day to warrant more than once per day?
There was a problem hiding this comment.
Once a day might actually be too much. Some managed services update the HO more than others but I do not think any of the them are in a place to do more than one update within a 24h period.
There was a problem hiding this comment.
Do you have any metrics on the velocity of code merged to HyperShift? The frequency of builds determines the size of the diff between nightlies, the larger the diff, the harder it will be to identify culprits of regression
do more than one update within a 24h period
I was under the impression that ARO stage was on a 4h delay from main presently. I know for prod they must wait 48h though
Some managed services update the HO more
There is a goal within managed services for everyone to be aligned on continuous delivery so I expect that to change at some point
There was a problem hiding this comment.
HyperShift usually gets ~40 PRs a week IIRC from the weekly reports I was running
|
|
||
| ## Proposal | ||
|
|
||
| Add a parallel, gated promotion path alongside the existing auto-release. A nightly pipeline resolves the most recent HO image built by Konflux's push build pipeline (triggered on every merge to `main`) and tests it against platform-specific e2e suites. Only tested images are promoted to a verified repository. Each platform's promotion is independent — a failure on one does not block others. |
There was a problem hiding this comment.
Are there any retest mechanisms here should it fail? Or is it then a case of wait until the next day?
Having this per platform makes the concept of a "green nightly" more elusive, is tracking the failures and escalation something you plan when there are consecutive failures?
There was a problem hiding this comment.
It's wait until the next day but retest is something we plan to follow up on later. It was not seen as a must have for a MVP.
There was a problem hiding this comment.
What's the story for "this went red three days in a row"? Who is monitoring that, how are they notified, what's the playbook?
There was a problem hiding this comment.
This would be one of the top RITS priorities
There was a problem hiding this comment.
The stale promotion alert (default 3 days, configurable) fires to the Slack channel when no successful promotion has occurred. As @celebdor noted, this would be a top RITS priority — RITS owns monitoring and triage of consecutive failures. The playbook is in the Support Procedures section: check PipelineRun logs, re-trigger for transient failures, escalate to HyperShift dev for persistent test failures. Updated the doc to make RITS ownership and the consecutive-failure escalation path explicit.
AI-assisted response via Claude Code
| | Phase | Coverage | Ownership | | ||
| | ----- | -------- | --------- | | ||
| | Phase 1 (MVP) | Cluster lifecycle, NodePool scaling, one upgrade path | HCP team — required for completion | | ||
| | Phase 2 | Full CPO version matrix (every supported 4.y.z and 4.y.0) | HCP team — required for completion | |
There was a problem hiding this comment.
What does this actually mean? Is this "run CPO on lots of 4.Y management clusters" or "CPO can create lots of 4.Y workload clusters"
There was a problem hiding this comment.
There is some CPO testing being done outside this effort but those tests will be included in the promotion process of the image. @clebs could point you to that effort.
There was a problem hiding this comment.
This is the epic that tracks improvements for multi-platform testing:
https://redhat.atlassian.net/browse/CNTRLPLANE-1849
| | Task | Condition | Purpose | | ||
| | ---- | --------- | ------- | | ||
| | `create-release` | `run-e2e` succeeded | Create Konflux Release object | | ||
| | `notify-slack` | `run-e2e` failed | Send Slack webhook notification | |
There was a problem hiding this comment.
The text above correctly states that tests run in Prow, but this task table still describes an in-Tekton execution model. setup-test-env, deploy-ho, and cleanup don't exist in the implementation — Prow manages all test infrastructure.
The actual pipeline has 3 tasks + a finally block:
| Task | Purpose |
|---|---|
extract-image |
Parse Snapshot, extract HO image reference |
run-e2e |
Trigger Prow periodic jobs via gangway API, poll for results |
evaluate-results |
Aggregate results from multiple parallel jobs (AND logic — all must pass) |
Finally tasks:
| Task | Condition | Purpose |
|---|---|---|
create-release |
evaluate-results succeeded |
Create Konflux Release object |
Note: once we migrate to the IntegrationTestScenario pattern (ITS with disabled context + CronJob labeling Snapshots), create-release will be removed — the Integration Service handles Release creation automatically via auto-release on test pass.
evaluate-results is also not in the current proposal — it supports running N test suites in parallel and gating on the aggregate result.
Ref: Slack confirmation
There was a problem hiding this comment.
Done. Updated the pipeline task table to reflect the Prow-based execution model — 3 tasks (extract-image, run-e2e via gangway, evaluate-results) plus notify-slack in the finally block. Removed setup-test-env, deploy-ho, and cleanup since Prow manages all test infrastructure.
AI-assisted response via Claude Code
|
|
||
| A new Tekton pipeline at `.tekton/pipelines/ho-release-gate.yaml` consisting of five sequential tasks followed by a `finally` block for promotion or notification. All test jobs run in Prow; Konflux launches Prow jobs and consumes their pass/fail results and run links. | ||
|
|
||
| **Pipeline tasks (sequential):** |
There was a problem hiding this comment.
With the IntegrationTestScenario pattern, Konflux offers a more idiomatic alternative to programmatic Release creation. Instead of a create-release finally task, the flow would be:
- Integration Service marks the Snapshot as passed/failed based on pipeline exit code
- ReleasePlan with
auto-release: "true"triggers Release creation automatically on pass - The Release is guaranteed to reference the exact tested Snapshot
This would remove the need for custom Release-creation logic in the pipeline and the gracePeriodDays field, and would also simplify the CR Interaction diagram above (the PR --> Release edge would go through the Integration Service instead).
Should we consider updating this section and the ReleasePlan definition to use auto-release instead of the explicit finally task?
There was a problem hiding this comment.
Done. Switched to the ITS auto-release pattern — ReleasePlan now uses auto-release: "true", removed the create-release finally task and the explicit Release YAML example, and updated the CR interaction diagram and sequence diagram to show the Integration Service handling Release creation on test pass.
AI-assisted response via Claude Code
|
|
||
| ### Non-Goals | ||
|
|
||
| 1. Replacing the existing auto-release to ACMD. The current auto-release path remains untouched. |
There was a problem hiding this comment.
Just a reminder — OCPSTRAT-3250 still says "Stop automatic release of HO images to ACMD. Images must pass platform-specific tests before being promoted." The enhancement proposal explicitly decided against this: auto-release to ACMD continues unchanged, and the gated pipeline is a parallel, additive path (not a replacement). We should update the OCPSTRAT description to reflect the enhancement proposal decisions.
There was a problem hiding this comment.
Good catch — the OCPSTRAT-3250 description predates the enhancement proposal's decision to keep auto-release unchanged. Will update the Jira description to reflect the parallel-path approach documented here.
AI-assisted response via Claude Code
Shows how CronJob, Snapshot, IntegrationTestScenario, PipelineRun, Release, ReleasePlan, and ReleasePlanAdmission interact during the nightly gating flow. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
@bryan-cox: all tests passed! Full PR test history. Your PR dashboard. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here. |
|
/LGTM |
Summary
OCPSTRAT-3250 / CNTRLPLANE-3434
Test plan
🤖 Generated with Claude Code