Skip to content

OCPSTRAT-3250: Konflux release gating pipeline for HyperShift Operator - #2016

Merged
openshift-merge-bot[bot] merged 3 commits into
openshift:masterfrom
bryan-cox:OCPSTRAT-3250
Jun 11, 2026
Merged

openshift-merge-bot[bot] merged 3 commits into
openshift:masterfrom
bryan-cox:OCPSTRAT-3250

Conversation

@bryan-cox

@bryan-cox bryan-cox commented May 19, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Adds a new enhancement proposal for a nightly Konflux-based release gating pipeline that validates HyperShift Operator images against e2e test suites before promoting them to verified repositories.
  • The pipeline operates alongside the existing auto-release to ACMD, adding a parallel, per-platform promotion path (ARO HCP pilot, ROSA HCP and GCP HCP future).
  • Includes full Konflux resource definitions (CronJob, ReleasePlan, IntegrationTestScenario, RBAC, Release), e2e test pipeline structure, error handling, strategy alignment, and related Jira issue tracking.

OCPSTRAT-3250 / CNTRLPLANE-3434

Test plan

  • markdownlint passes cleanly
  • All required enhancement template headings present
  • YAML frontmatter validates

🤖 Generated with Claude Code

@openshift-ci
openshift-ci Bot requested review from csrwng and enxebre May 19, 2026 17:26
@bryan-cox bryan-cox changed the title Enhancement: Konflux release gating pipeline for HyperShift Operator OCPSTRAT-3250: Enhancement: Konflux release gating pipeline for HyperShift Operator May 19, 2026
@openshift-ci-robot openshift-ci-robot added the jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. label May 19, 2026
@openshift-ci-robot

openshift-ci-robot commented May 19, 2026 •

Copy link
Copy Markdown

@bryan-cox: This pull request references OCPSTRAT-3250 which is a valid jira issue.

Details

In response to this:

Summary

  • Adds a new enhancement proposal for a nightly Konflux-based release gating pipeline that validates HyperShift Operator images against e2e test suites before promoting them to verified repositories.
  • The pipeline operates alongside the existing auto-release to ACMD, adding a parallel, per-platform promotion path (ARO HCP pilot, ROSA HCP and GCP HCP future).
  • Includes full Konflux resource definitions (CronJob, ReleasePlan, IntegrationTestScenario, RBAC, Release), e2e test pipeline structure, error handling, strategy alignment, and related Jira issue tracking.

OCPSTRAT-3250 / CNTRLPLANE-3434

Test plan

  • markdownlint passes cleanly
  • All required enhancement template headings present
  • YAML frontmatter validates

🤖 Generated with Claude Code

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@bryan-cox bryan-cox changed the title OCPSTRAT-3250: Enhancement: Konflux release gating pipeline for HyperShift Operator OCPSTRAT-3250: Konflux release gating pipeline for HyperShift Operator May 19, 2026
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
@bryan-cox
bryan-cox force-pushed the OCPSTRAT-3250 branch 2 times, most recently from e8dd6e1 to 68fbb5f Compare May 19, 2026 18:05
@bryan-cox

Copy link
Copy Markdown
Member Author

While we are implementing this effort for ARO HCP first, we are expecting to onboard ROSA HCP and GCP in the future. I wanted to make sure y'all were aware of this enhancement; please feel free to unsubscribe if you wish - @deads2k @joshbranham @cblecker

@bryan-cox
bryan-cox force-pushed the OCPSTRAT-3250 branch 2 times, most recently from 5f6b081 to a135a21 Compare May 20, 2026 14:39
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Introduces a nightly, platform-independent gating system that validates
HyperShift Operator images against e2e test suites before promoting them
to verified repositories. The pipeline operates alongside the existing
Konflux auto-release, adding a parallel promotion path per managed service
platform (ARO HCP pilot, ROSA HCP and GCP HCP future).

OCPSTRAT-3250

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md Outdated
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
Comment thread enhancements/hypershift/konflux-release-gating-pipeline.md
@enxebre

enxebre commented May 28, 2026

Copy link
Copy Markdown
Member

dropped some more questions, lgtm

- Add Design Rationale section capturing nightly cadence and alerting
  rationale from PR discussion threads (enxebre)
- Clarify per-platform pipeline naming strategy with TBD for shared vs
  separate files when second platform onboards (enxebre)
- Add Verified Repository to glossary as single shared quay.io repo
  with per-platform image tags (enxebre)
- Document resource management: pipeline in .tekton/pipelines/,
  Konflux namespace resources in contrib/konflux/ (enxebre)
- Fix CronJob/ITS interaction to follow Konflux periodic test pattern:
  ITS uses disabled context, CronJob labels snapshots instead of
  creating PipelineRuns directly (Nirshal)
- Update RBAC, diagrams, and workflow steps to match new pattern

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@enxebre

enxebre commented May 28, 2026

Copy link
Copy Markdown
Member

/approve

@openshift-ci

openshift-ci Bot commented May 28, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: enxebre

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@openshift-ci openshift-ci Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label May 28, 2026

@JoelSpeed JoelSpeed left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One thing I'm hoping we solve as part of this EP is the ability to have confidence in the supported matrix of CPO to guest and CPO to management version skews. There is only a very light mention of cross version testing here, was that something you were considering in/out of scope?


## Summary

This enhancement introduces a nightly, platform-independent gating system that validates HyperShift Operator (HO) images against end-to-end test suites before promoting them to verified repositories. The pipeline operates alongside the existing Konflux auto-release mechanism, adding a parallel promotion path that only advances images which have passed e2e test suites agreed upon between the HyperShift and managed service (HCM) teams. Tests may vary by platform.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Promotion requires all tests pass across all platforms? Or are there separate promotion destinations such that we might see promotion succeed on ARO but not ROSA?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There will be different promotion paths for each managed service since each one needs a different set of tests to pass. If ARO HCP tests fail but GCP and ROSA tests pass, they should still get a tagged HO for their managed services respectively.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How do you see that increasing complexity for the HyperShift team? You'd end up with a partially green nightly rather than a full green or full red. It brings in complexity like having to understand "ok this PR went out to GCP and AWS but not Azure because XYZ"

I wonder if it's desirable from the HyperShift team perspective to just have a single signal for green vs red?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the benefit of not blocking releases for issues that affect a single managed platform justifies having split reporting.


#### Design Rationale

**Nightly cadence (24h):** Each pipeline run provisions real cloud infrastructure with platform-specific credentials (e.g., Azure for ARO HCP). A nightly cadence balances validation confidence with cloud infrastructure cost. Per-commit gating is cost-prohibitive and would slow the development feedback loop (see Alternatives). Per-platform cadence can differ — each platform can have its own CronJob schedule.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OpenShift CI and nightly builds happen every 6h, have you considered making this more frequent than once per day? Is there enough change in a day to warrant more than once per day?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Once a day might actually be too much. Some managed services update the HO more than others but I do not think any of the them are in a place to do more than one update within a 24h period.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have any metrics on the velocity of code merged to HyperShift? The frequency of builds determines the size of the diff between nightlies, the larger the diff, the harder it will be to identify culprits of regression

do more than one update within a 24h period

I was under the impression that ARO stage was on a 4h delay from main presently. I know for prod they must wait 48h though

Some managed services update the HO more

There is a goal within managed services for everyone to be aligned on continuous delivery so I expect that to change at some point

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

HyperShift usually gets ~40 PRs a week IIRC from the weekly reports I was running


## Proposal

Add a parallel, gated promotion path alongside the existing auto-release. A nightly pipeline resolves the most recent HO image built by Konflux's push build pipeline (triggered on every merge to `main`) and tests it against platform-specific e2e suites. Only tested images are promoted to a verified repository. Each platform's promotion is independent — a failure on one does not block others.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there any retest mechanisms here should it fail? Or is it then a case of wait until the next day?

Having this per platform makes the concept of a "green nightly" more elusive, is tracking the failures and escalation something you plan when there are consecutive failures?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's wait until the next day but retest is something we plan to follow up on later. It was not seen as a must have for a MVP.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the story for "this went red three days in a row"? Who is monitoring that, how are they notified, what's the playbook?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would be one of the top RITS priorities

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The stale promotion alert (default 3 days, configurable) fires to the Slack channel when no successful promotion has occurred. As @celebdor noted, this would be a top RITS priority — RITS owns monitoring and triage of consecutive failures. The playbook is in the Support Procedures section: check PipelineRun logs, re-trigger for transient failures, escalate to HyperShift dev for persistent test failures. Updated the doc to make RITS ownership and the consecutive-failure escalation path explicit.


AI-assisted response via Claude Code

| Phase | Coverage | Ownership |
| ----- | -------- | --------- |
| Phase 1 (MVP) | Cluster lifecycle, NodePool scaling, one upgrade path | HCP team — required for completion |
| Phase 2 | Full CPO version matrix (every supported 4.y.z and 4.y.0) | HCP team — required for completion |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does this actually mean? Is this "run CPO on lots of 4.Y management clusters" or "CPO can create lots of 4.Y workload clusters"

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is some CPO testing being done outside this effort but those tests will be included in the promotion process of the image. @clebs could point you to that effort.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the epic that tracks improvements for multi-platform testing:
https://redhat.atlassian.net/browse/CNTRLPLANE-1849

| Task | Condition | Purpose |
| ---- | --------- | ------- |
| `create-release` | `run-e2e` succeeded | Create Konflux Release object |
| `notify-slack` | `run-e2e` failed | Send Slack webhook notification |

@Nirshal Nirshal Jun 11, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The text above correctly states that tests run in Prow, but this task table still describes an in-Tekton execution model. setup-test-env, deploy-ho, and cleanup don't exist in the implementation — Prow manages all test infrastructure.

The actual pipeline has 3 tasks + a finally block:

Task Purpose
extract-image Parse Snapshot, extract HO image reference
run-e2e Trigger Prow periodic jobs via gangway API, poll for results
evaluate-results Aggregate results from multiple parallel jobs (AND logic — all must pass)

Finally tasks:

Task Condition Purpose
create-release evaluate-results succeeded Create Konflux Release object

Note: once we migrate to the IntegrationTestScenario pattern (ITS with disabled context + CronJob labeling Snapshots), create-release will be removed — the Integration Service handles Release creation automatically via auto-release on test pass.

evaluate-results is also not in the current proposal — it supports running N test suites in parallel and gating on the aggregate result.

Ref: Slack confirmation

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Updated the pipeline task table to reflect the Prow-based execution model — 3 tasks (extract-image, run-e2e via gangway, evaluate-results) plus notify-slack in the finally block. Removed setup-test-env, deploy-ho, and cleanup since Prow manages all test infrastructure.


AI-assisted response via Claude Code


A new Tekton pipeline at `.tekton/pipelines/ho-release-gate.yaml` consisting of five sequential tasks followed by a `finally` block for promotion or notification. All test jobs run in Prow; Konflux launches Prow jobs and consumes their pass/fail results and run links.

**Pipeline tasks (sequential):**

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the IntegrationTestScenario pattern, Konflux offers a more idiomatic alternative to programmatic Release creation. Instead of a create-release finally task, the flow would be:

  1. Integration Service marks the Snapshot as passed/failed based on pipeline exit code
  2. ReleasePlan with auto-release: "true" triggers Release creation automatically on pass
  3. The Release is guaranteed to reference the exact tested Snapshot

This would remove the need for custom Release-creation logic in the pipeline and the gracePeriodDays field, and would also simplify the CR Interaction diagram above (the PR --> Release edge would go through the Integration Service instead).

Should we consider updating this section and the ReleasePlan definition to use auto-release instead of the explicit finally task?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Switched to the ITS auto-release pattern — ReleasePlan now uses auto-release: "true", removed the create-release finally task and the explicit Release YAML example, and updated the CR interaction diagram and sequence diagram to show the Integration Service handling Release creation on test pass.


AI-assisted response via Claude Code


### Non-Goals

1. Replacing the existing auto-release to ACMD. The current auto-release path remains untouched.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a reminder — OCPSTRAT-3250 still says "Stop automatic release of HO images to ACMD. Images must pass platform-specific tests before being promoted." The enhancement proposal explicitly decided against this: auto-release to ACMD continues unchanged, and the gated pipeline is a parallel, additive path (not a replacement). We should update the OCPSTRAT description to reflect the enhancement proposal decisions.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — the OCPSTRAT-3250 description predates the enhancement proposal's decision to keep auto-release unchanged. Will update the Jira description to reflect the parallel-path approach documented here.


AI-assisted response via Claude Code

Shows how CronJob, Snapshot, IntegrationTestScenario, PipelineRun,
Release, ReleasePlan, and ReleasePlanAdmission interact during
the nightly gating flow.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@openshift-ci

openshift-ci Bot commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

@bryan-cox: all tests passed!

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@davegord

Copy link
Copy Markdown

/LGTM

@openshift-ci openshift-ci Bot added the lgtm Indicates that a PR is ready to be merged. label Jun 11, 2026
@openshift-merge-bot
openshift-merge-bot Bot merged commit 064ac37 into openshift:master Jun 11, 2026
2 checks passed
@bryan-cox
bryan-cox deleted the OCPSTRAT-3250 branch June 17, 2026 15:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. lgtm Indicates that a PR is ready to be merged.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants