Skip to content

CNTRLPLANE-3434: add ho-release-gate pipeline for nightly promotion - #8602

Merged
openshift-merge-bot[bot] merged 8 commits into
openshift:mainfrom
Nirshal:ho-release-gate-pipeline
Jul 21, 2026
Merged

openshift-merge-bot[bot] merged 8 commits into
openshift:mainfrom
Nirshal:ho-release-gate-pipeline

Conversation

@Nirshal

@Nirshal Nirshal commented May 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds the Tekton pipeline and PipelineRun template for the HyperShift Operator nightly release gating pipeline, triggered via IntegrationTestScenario (ITS).

Pipeline flow

CronJob (3:15 UTC nightly)
  |-- Resolves latest Konflux Snapshot (auto-released, push build)
  |-- Labels Snapshot with test.appstudio.openshift.io/run=<its-name>

Integration Service
  |-- Detects labeled Snapshot, creates PipelineRun from template
  |-- Injects SNAPSHOT JSON payload + ITS params into PipelineRun

PipelineRun (ho-release-gate):
  extract-image     - Parses SNAPSHOT JSON to extract HO image and commit SHA.
                      Gets snapshot name from PipelineRun label.
  run-e2e           - Triggers Prow periodic e2e jobs via gangway API.
                      Two categories: blocking (gate fails if any fail) and
                      informing (reported but do not block the gate).
                      45min initial delay, then polls every 10 min with 30s stagger.
                      Overrides: MULTISTAGE_PARAM_OVERRIDE_OVERRIDE_HYPERSHIFT_OPERATOR_IMAGE
                                 (bypasses ci-operator ImageStream race condition),
                                 OVERRIDE_IMAGE_HYPERSHIFT_TESTS.
                      Always exits 0 - writes per-job results JSON (with type field).
  evaluate-results  - Reads results, applies AND logic on blocking tests only.
                      Informing failures logged as warnings.
                      Always exits 0 - writes gate-passed result (true/false).
  create-release    - If gate passed: creates Release CR referencing the ReleasePlan
                      and Snapshot, writes release name to result.
                      If gate failed: writes "N/A" to result, exits 1
                      (marks pipeline as Failed in UI).
  finally:
    notify-slack       - Fires on every run (pass or fail).
                         Three visual states:
                           Green (#2E7D32): gate passed, Release created.
                           Orange (#F57C00): gate passed, Release creation failed.
                           Red (#D32F2F): gate failed.
                         Gate-label in header, grouped results (blocking + informing).
                         On gate failure: checks for stale promotion streak via
                         KubeArchive. If streak >= threshold, sends stale alert
                         instead of normal failure notification.
    notify-slack-error - Fallback for catastrophic failures (DAG failure).
                         Zero result dependencies, fires when create-release
                         did not run. Gate-label in header.
                         Also checks for stale promotion streak via KubeArchive.
                         Sends stale alert or generic error notification.

Release (on gate pass):
  create-release task creates Release CR explicitly.
  ReleasePlan targets rhtap-releng-tenant (managed workspace).
  Managed pipeline pushes image to verified Quay repo.

Stale promotion alerting (CNTRLPLANE-3451)

When the gate fails, both notify-slack and notify-slack-error query the
KubeArchive REST API for historical PipelineRuns matching the ITS scenario label.
The consecutive failure streak is measured in days (difference between now and the
oldest consecutive failed run, inclusive of today).

If streak_days >= stale-threshold-days, the normal notification is replaced with
a stale promotion alert containing:

  • Header with gate label and streak duration
  • Fields: Gate, Threshold, Streak, Current PipelineRun
  • History of up to 10 most recent failed runs with clickable PipelineRun links,
    dates, and failure reasons (Failed, PipelineRunTimeout, CouldntGetPipeline, etc.)
  • Truncation summary for streaks longer than 10 runs
  • Footer with investigation guidance

The stale check is safe to skip: if KubeArchive is unreachable or returns no data,
the normal notification is sent instead.

Notification reliability

All tasks use default-first initialization: each task writes safe default values
to its result paths at the start of its script, before any logic. If a task starts
but crashes mid-execution (e.g., OOMKilled), results are still initialized.

Two mutually exclusive finally tasks ensure a notification is always sent:

  • notify-slack: detailed notification, depends on task results (skipped if results uninitialized)
  • notify-slack-error: generic fallback, zero result dependencies, fires only when create-release was skipped by DAG failure

Both finally tasks perform the stale promotion check independently, so a stale alert
is sent regardless of which notification path fires.

Files

File Description
.tekton/pipelines/ho-release-gate.yaml Pipeline definition
.tekton/pipelines/ho-release-gate-run.yaml PipelineRun template (referenced by ITS via git resolver)
.tekton/lib/ho_release_gate.py Pipeline orchestration: image extraction, job triggering/polling, gate evaluation, notifications, stale alerting
.tekton/lib/kubearchive_utils.py KubeArchive REST API: fetch historical PipelineRuns
.tekton/lib/prow_utils.py Gangway API: trigger jobs, resolve URLs, poll status
.tekton/lib/slack_utils.py Slack webhook: send messages, build Block Kit payloads
.tekton/lib/http_utils.py Low-level HTTP wrapper with retry logic (stdlib only)
.tekton/lib/README.md Library documentation: architecture, constraints, runtime environment
.tekton/lib/tests/ Unit tests for all modules (unittest + unittest.mock, stdlib only)
.tekton/lib/mock/ Mock utilities for pipeline integration testing

Pipeline parameters

Parameter Source Description
SNAPSHOT Integration Service (auto-injected) JSON payload of the Konflux Snapshot
e2e-blocking-job-names ITS params JSON array of blocking Prow periodic job names
e2e-informing-job-names ITS params JSON array of informing Prow periodic job names
gate-label ITS params Label identifying which service gate (e.g. ARO HCP)
release-plan-name ITS params Name of the ReleasePlan to reference when creating the Release CR
stale-threshold-days ITS params Number of consecutive failure days before triggering stale alert (default: 3)

Resources

All resources live in namespace crt-redhat-acm-tenant unless noted.

Resource Name Managed via
Pipeline + PipelineRun template .tekton/pipelines/ho-release-gate*.yaml This PR
ReleasePlanAdmission redhat-hypershift-operator-ho-release-gate-aro-hcp (in rhtap-releng-tenant) GitOps (!18934 - Merged)
ReleasePlan hypershift-operator-ho-release-gate-aro-hcp GitOps (!18938 - Merged)
CronJob hypershift-operator-nightly-promotion GitOps (!18938 - Merged)
IntegrationTestScenario hypershift-ho-release-gate-aro-hcp GitOps (!19261 - Merged)
ServiceAccount nightly-promotion-sa GitOps (!18934 - Merged)
RBAC (snapshot labeling) nightly-promotion-sa -> konflux-tester-internalbot-actions GitOps (!19870 - Merged)
RBAC (release creation) nightly-promotion-sa -> konflux-releaser-bot-actions via taskRunSpecs GitOps (!19870 - Merged)
Secret gangway-token (Prow auth) Manual (not in GitOps)
Secret slack-webhook (Slack notifications) Manual (not in GitOps)

Release mechanism

The pipeline uses explicit Release CR creation (not auto-release):

  1. The create-release task creates a Release object referencing the ReleasePlan (name passed via ITS param release-plan-name) and the tested Snapshot
  2. The ReleasePlan targets rhtap-releng-tenant (managed workspace)
  3. The RPA (redhat-hypershift-operator-ho-release-gate-aro-hcp) picks up the Release
  4. The managed pipeline (rh-push-to-external-registry) pushes the validated image to the verified Quay repo with platform-prefixed tags (aro-hcp-latest, aro-hcp-latest-{{ timestamp }}, aro-hcp-{{ git_sha }}, etc.)

Target repo: quay.io/redhat-services-prod/crt-redhat-acm-tenant/hypershift/hypershift-operator-verified

Note: Auto-release was disabled (!19296 - Merged) because with contexts: disabled on the ITS, Integration Service was auto-releasing every normal push snapshot through our ReleasePlan, bypassing the release gate entirely.

Remaining work

Item Status Notes
HO image override fix Blocked Depends on openshift/release#81877 (adds OVERRIDE_HYPERSHIFT_OPERATOR_IMAGE step parameter + Gangway transport variable to hypershift-install)
Pipeline source migration On merge ITS resolverRef will point to openshift/hypershift:main once this PR merges (!19913 - Draft)
Regression analysis Future Component-readiness tracking (deads2k requirement)

Related

Summary by CodeRabbit

  • New Features
    • Added a new Tekton release gate pipeline that parses the provided snapshot, runs configured Gangway Prow jobs, and computes a gate verdict.
    • Release creation is skipped when blocking checks don’t pass.
    • Added Slack notifications for normal results, errors, and stale-promotion conditions; notifications run even on failures.
  • Documentation
    • Documented the shared stdlib-only release-gate helper library and how to run its tests.
  • Tests
    • Added unit test coverage for HTTP retries, Prow/Gangway helpers, gate evaluation, Slack payloads, and stale detection logic.

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Pipeline controller notification
This repo is configured to use the pipeline controller. Second-stage tests will be triggered either automatically or after lgtm label is added, depending on the repository configuration. The pipeline controller will automatically detect which contexts are required and will utilize /test Prow commands to trigger the second stage.

For optional jobs, comment /test ? to see a list of all defined jobs. To trigger manually all jobs from second stage use /pipeline required command.

This repository is configured in: LGTM mode

@coderabbitai

coderabbitai Bot commented May 27, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 60d72f07-60b5-4afd-920e-51e6ce22f011

📥 Commits

Reviewing files that changed from the base of the PR and between 8dedf95 and 7194160.

📒 Files selected for processing (6)
  • .tekton/lib/ho_release_gate.py
  • .tekton/lib/kubearchive_utils.py
  • .tekton/lib/tests/test_ho_release_gate.py
  • .tekton/lib/tests/test_kubearchive_utils.py
  • .tekton/pipelines/ho-release-gate-run.yaml
  • .tekton/pipelines/ho-release-gate.yaml
🚧 Files skipped from review as they are similar to previous changes (4)
  • .tekton/pipelines/ho-release-gate-run.yaml
  • .tekton/lib/tests/test_kubearchive_utils.py
  • .tekton/lib/ho_release_gate.py
  • .tekton/lib/tests/test_ho_release_gate.py

📝 Walkthrough

Walkthrough

The PR adds a Tekton HO release gate that extracts a Snapshot image, runs blocking and informing Gangway jobs, evaluates their results, conditionally creates a Konflux Release, and sends Slack notifications. Shared stdlib-only helpers provide HTTP, Prow, KubeArchive, Slack, orchestration, stale-history handling, mocks, documentation, and unit tests. A PipelineRun configures timeouts, workspace storage, service-account usage, and pipeline resolution.

Sequence Diagram(s)

sequenceDiagram
  participant PipelineRun
  participant TektonPipeline
  participant Gangway
  participant Konflux
  participant Slack

  PipelineRun->>TektonPipeline: provide Snapshot and gate parameters
  TektonPipeline->>Gangway: trigger and poll E2E jobs
  Gangway-->>TektonPipeline: return job outcomes
  TektonPipeline->>TektonPipeline: evaluate gate
  TektonPipeline->>Konflux: create Release when gate passes
  TektonPipeline->>Slack: send gate or error notification
Loading

Suggested reviewers: sjenning, enxebre


Important

Pre-merge checks failed

Please resolve all errors before merging. Addressing warnings is optional.

❌ Failed checks (1 error)

Check name Status Explanation Resolution
No-Sensitive-Data-In-Logs ❌ Error FAIL: The pipeline logs per-job URLs and raw results JSON, which can leak internal CI hostnames in stdout. Redact or remove URL/raw-body prints; log only job names/statuses and avoid echoing any bearer-token-derived or API-response content.
✅ Passed checks (10 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title matches the main change by describing the new ho-release-gate CI pipeline for nightly promotion.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Stable And Deterministic Test Names ✅ Passed No changed Ginkgo tests were found; the added tests are Python unittest methods with static names only.
Test Structure And Quality ✅ Passed The PR adds only Python unittest tests under .tekton/lib/tests; there are no Ginkgo It blocks or cluster-interacting tests, so this check is not applicable.
Topology-Aware Scheduling Compatibility ✅ Passed PASS: The PR only adds Tekton pipeline/task scripts and tests; no nodeSelector, affinity, topologySpreadConstraints, replicas, or PDBs were introduced.
Ipv6 And Disconnected Network Test Compatibility ✅ Passed No new Ginkgo e2e specs were added; the branch changes are Tekton YAML, Python unit tests/helpers, and Go helper/unit-test code only.
No-Weak-Crypto ✅ Passed No MD5/SHA1/DES/RC4/3DES/Blowfish/ECB, custom crypto, or secret/token comparisons were found in the touched .tekton files.
Container-Privileges ✅ Passed No privileged fields, host namespace flags, SYS_ADMIN, or explicit root/allowPrivilegeEscalation settings appear in the new Tekton manifests.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@openshift-ci openshift-ci Bot added do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. do-not-merge/needs-area labels May 27, 2026
@openshift-ci

openshift-ci Bot commented May 27, 2026

Copy link
Copy Markdown
Contributor

Skipping CI for Draft Pull Request.
If you want CI signal for your change, please convert it to an actual PR.
You can still manually trigger a test run with /test all

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
.tekton/pipelines/ho-release-gate.yaml (2)

91-102: ⚡ Quick win

Consider adding a timeout to the polling loop.

The commented implementation polls indefinitely until success/failure. If the Prow job gets stuck or the API becomes unreachable, this could cause the pipeline to hang forever.

When implementing the actual gangway integration, add a maximum retry count or deadline:

♻️ Suggested pattern for timeout
+          MAX_ATTEMPTS=180  # 3 hours at 60s intervals
+          ATTEMPT=0
           # Poll for completion:
           while true; do
+            ATTEMPT=$((ATTEMPT + 1))
+            if [[ ${ATTEMPT} -gt ${MAX_ATTEMPTS} ]]; then
+              echo "ERROR: Timeout waiting for Prow job completion"
+              echo -n "failed" > $(results.result.path)
+              break
+            fi
             STATUS=$(curl -s "${GANGWAY_URL}/v1/executions/${JOB_URL}" \
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.tekton/pipelines/ho-release-gate.yaml around lines 91 - 102, The commented
polling loop for checking Gangway job status (using GANGWAY_URL, JOB_URL,
GANGWAY_TOKEN and writing to results.result.path) can hang indefinitely; update
the loop to enforce a timeout by adding either a max retry counter or a deadline
variable (e.g., MAX_RETRIES or GANGWAY_POLL_DEADLINE_SECONDS) and break with a
failure result when exceeded; ensure the loop increments the counter or checks
the deadline each iteration, logs a clear timeout error, and writes "failed" to
results.result.path if the timeout is reached.

29-29: 💤 Low value

Consider pinning container image versions for reproducibility.

Multiple tasks use :latest tags (lines 29, 62, 124, 164). For a release gate pipeline, unexpected image updates could cause inconsistent behavior or breakages. Pin to specific digests or version tags before removing the draft status.

♻️ Example with pinned versions
-        image: registry.redhat.io/openshift4/ose-cli:latest
+        image: registry.redhat.io/openshift4/ose-cli:v4.15

Or use digest for stronger guarantees:

image: registry.redhat.io/openshift4/ose-cli@sha256:<digest>
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.tekton/pipelines/ho-release-gate.yaml at line 29, Replace occurrences of
the image field using the :latest tag (e.g.,
"registry.redhat.io/openshift4/ose-cli:latest") with explicit, pinned version
tags or immutable digests (e.g., "`@sha256`:...") to ensure reproducible builds;
update every task that references the same image (the other occurrences of the
same "image: registry.redhat.io/openshift4/ose-cli:latest" in this pipeline) to
the chosen tag/digest and verify compatibility before merging.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.tekton/pipelines/ho-release-gate.yaml:
- Around line 112-153: The pipeline currently only handles explicit
$(tasks.run-e2e.results.result) values "passed" or "failed" so errors/skips
produce no notification; add a catch-all task (e.g., notify-error) or extend
notify-slack to inspect $(tasks.run-e2e.status) so non-Succeeded statuses
trigger a notification. Specifically, add a finally task (name: notify-error)
using when: input: $(tasks.run-e2e.status) operator: notin values: ["Succeeded"]
(and/or guard with $(tasks.run-e2e.results.result) notin ["passed","failed"])
that sends an alert, or update the existing notify-slack when clause to include
$(tasks.run-e2e.status) operator: in values: ["Failed"] so task
errors/timeouts/omissions are reported.
- Around line 165-171: The Slack JSON payload is built by interpolating params
directly in the shell script (the script block that posts to
"${SLACK_WEBHOOK_URL}"), which risks JSON injection if params like
$(params.ho-image), $(params.snapshot-name) or $(params.prow-job-url) contain
quotes or newlines; fix by constructing the JSON safely with a JSON tool (e.g.,
use jq -n --arg snapshot "$(params.snapshot-name)" --arg image
"$(params.ho-image)" --arg prow "$(params.prow-job-url)" '{text: "HyperShift
nightly promotion FAILED\nSnapshot: \($snapshot)\nImage: \($image)\nProw job:
\($prow)\nPipeline: $(context.pipelineRun.name)"}' ) so each param is properly
escaped and then pipe that output to curl instead of embedding parameters
directly in the here-doc.

---

Nitpick comments:
In @.tekton/pipelines/ho-release-gate.yaml:
- Around line 91-102: The commented polling loop for checking Gangway job status
(using GANGWAY_URL, JOB_URL, GANGWAY_TOKEN and writing to results.result.path)
can hang indefinitely; update the loop to enforce a timeout by adding either a
max retry counter or a deadline variable (e.g., MAX_RETRIES or
GANGWAY_POLL_DEADLINE_SECONDS) and break with a failure result when exceeded;
ensure the loop increments the counter or checks the deadline each iteration,
logs a clear timeout error, and writes "failed" to results.result.path if the
timeout is reached.
- Line 29: Replace occurrences of the image field using the :latest tag (e.g.,
"registry.redhat.io/openshift4/ose-cli:latest") with explicit, pinned version
tags or immutable digests (e.g., "`@sha256`:...") to ensure reproducible builds;
update every task that references the same image (the other occurrences of the
same "image: registry.redhat.io/openshift4/ose-cli:latest" in this pipeline) to
the chosen tag/digest and verify compatibility before merging.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: bd039c9f-b76d-4301-b332-b5f6622f4d01

📥 Commits

Reviewing files that changed from the base of the PR and between 09c7701 and 56fc5a7.

📒 Files selected for processing (1)
  • .tekton/pipelines/ho-release-gate.yaml

Comment thread .tekton/pipelines/ho-release-gate.yaml Outdated
Comment thread .tekton/pipelines/ho-release-gate.yaml Outdated
@codecov

codecov Bot commented May 27, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 44.12%. Comparing base (433a404) to head (7194160).
⚠️ Report is 150 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #8602      +/-   ##
==========================================
+ Coverage   43.60%   44.12%   +0.51%     
==========================================
  Files         771      779       +8     
  Lines       95806    98065    +2259     
==========================================
+ Hits        41778    43269    +1491     
- Misses      51119    51761     +642     
- Partials     2909     3035     +126     

see 66 files with indirect coverage changes

Flag Coverage Δ
cmd-support 38.24% <ø> (+1.01%) ⬆️
cpo-hostedcontrolplane 46.27% <ø> (+0.35%) ⬆️
cpo-other 45.22% <ø> (+0.11%) ⬆️
hypershift-operator 54.14% <ø> (+0.48%) ⬆️
other 33.59% <ø> (+1.50%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
.tekton/pipelines/ho-release-gate.yaml (1)

230-234: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Harden Slack webhook call with fail-fast and timeout controls.

The webhook POST can currently fail silently (non-2xx) or hang without bounds. Add curl failure/timeout/retry flags so pipeline outcome reflects notification delivery failures.

Proposed patch
       - name: send-notification
         image: curlimages/curl:latest
         script: |
           #!/bin/sh
-          curl -X POST -H 'Content-type: application/json' \
+          curl --fail --show-error --silent \
+            --connect-timeout 10 --max-time 30 \
+            --retry 3 --retry-delay 2 \
+            -X POST -H 'Content-type: application/json' \
             --data "{
               \"text\": \"HyperShift nightly promotion FAILED\nSnapshot: $(params.snapshot-name)\nImage: $(params.ho-image)\nProw job: $(params.prow-job-url)\nPipeline: $(context.pipelineRun.name)\"
             }" \
             "${SLACK_WEBHOOK_URL}"
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.tekton/pipelines/ho-release-gate.yaml around lines 230 - 234, Update the
curl invocation that posts to "${SLACK_WEBHOOK_URL}" (the Slack webhook step) to
use robust failure, timeout, and retry flags so non-2xx responses and hangs
produce a non-zero exit: add --fail --show-error --connect-timeout 5 --max-time
10 --retry 3 --retry-delay 2 --retry-connrefused to the existing curl command
that posts the JSON payload (the block using params.snapshot-name,
params.ho-image, params.prow-job-url and context.pipelineRun.name) so the
pipeline reflects notification delivery failures.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In @.tekton/pipelines/ho-release-gate.yaml:
- Around line 230-234: Update the curl invocation that posts to
"${SLACK_WEBHOOK_URL}" (the Slack webhook step) to use robust failure, timeout,
and retry flags so non-2xx responses and hangs produce a non-zero exit: add
--fail --show-error --connect-timeout 5 --max-time 10 --retry 3 --retry-delay 2
--retry-connrefused to the existing curl command that posts the JSON payload
(the block using params.snapshot-name, params.ho-image, params.prow-job-url and
context.pipelineRun.name) so the pipeline reflects notification delivery
failures.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 0389e56d-fe3e-4fce-a4a9-7608f8e44e1b

📥 Commits

Reviewing files that changed from the base of the PR and between 56fc5a7 and d079b12.

📒 Files selected for processing (1)
  • .tekton/pipelines/ho-release-gate.yaml

@Nirshal
Nirshal force-pushed the ho-release-gate-pipeline branch 2 times, most recently from e79809d to 58a3243 Compare May 28, 2026 15:45
@Nirshal Nirshal changed the title WIP: Add ho-release-gate pipeline for nightly promotion feat(ci): add ho-release-gate pipeline for nightly promotion May 28, 2026
@coderabbitai

coderabbitai Bot commented May 28, 2026

Copy link
Copy Markdown
Contributor

Actionable comments posted: 0

@Nirshal
Nirshal force-pushed the ho-release-gate-pipeline branch from 58a3243 to bfb44b0 Compare May 28, 2026 16:31

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
.tekton/pipelines/ho-release-gate.yaml (1)

112-118: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Add a non-success notification path in finally.

Only the pass path is handled today. If run-e2e returns failed or errors before publishing results, there is no explicit failure notification task in this pipeline. Add a failure/error finalizer path keyed off task status/results.

💡 Minimal fix pattern
   finally:
   - name: create-release
     when:
     - input: $(tasks.run-e2e.results.result)
       operator: in
       values: ["passed"]
@@
     - name: snapshot-name
       value: $(tasks.extract-image.results.snapshot-name)
+
+  - name: notify-failure
+    when:
+    - input: $(tasks.run-e2e.status)
+      operator: notin
+      values: ["Succeeded"]
+    taskSpec:
+      steps:
+      - name: notify
+        image: registry.redhat.io/openshift4/ose-cli:latest
+        script: |
+          #!/bin/bash
+          set -euo pipefail
+          echo "run-e2e did not complete successfully; add Slack notification here"
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.tekton/pipelines/ho-release-gate.yaml around lines 112 - 118, The
pipeline's finally currently only handles the success path for the
create-release finalizer (referencing the create-release entry and the run-e2e
task via $(tasks.run-e2e.results.result)); add a complementary finalizer (e.g.,
name: notify-failure) that triggers when run-e2e did not pass by using a when
clause such as input: $(tasks.run-e2e.results.result) operator: notin values:
["passed"] and also add a guard on $(tasks.run-e2e.status) to catch missing
results/errors (e.g., operator: in values: ["Failed","Error"] or similar), and
implement the notification taskSpec for failure/error handling; ensure both
create-release and notify-failure entries live under finally so failures are
explicitly handled.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.tekton/pipelines/ho-release-gate.yaml:
- Around line 36-39: The pipeline currently validates params.ho-image but does
not check params.snapshot-name, so add an early non-empty validation for
snapshot-name in the same extract-image validation block: detect if
"$(params.snapshot-name)" is empty, emit a clear error like "ERROR:
snapshot-name parameter is empty" and exit 1 to fail fast; locate the validation
near the existing ho-image check in the extract-image task/script and mirror the
same pattern to ensure bad input fails before e2e runs.

---

Duplicate comments:
In @.tekton/pipelines/ho-release-gate.yaml:
- Around line 112-118: The pipeline's finally currently only handles the success
path for the create-release finalizer (referencing the create-release entry and
the run-e2e task via $(tasks.run-e2e.results.result)); add a complementary
finalizer (e.g., name: notify-failure) that triggers when run-e2e did not pass
by using a when clause such as input: $(tasks.run-e2e.results.result) operator:
notin values: ["passed"] and also add a guard on $(tasks.run-e2e.status) to
catch missing results/errors (e.g., operator: in values: ["Failed","Error"] or
similar), and implement the notification taskSpec for failure/error handling;
ensure both create-release and notify-failure entries live under finally so
failures are explicitly handled.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: abc02483-34a2-44b5-865b-99e67bdfd6e4

📥 Commits

Reviewing files that changed from the base of the PR and between 58a3243 and bfb44b0.

📒 Files selected for processing (1)
  • .tekton/pipelines/ho-release-gate.yaml

Comment thread .tekton/pipelines/ho-release-gate.yaml Outdated
@Nirshal
Nirshal force-pushed the ho-release-gate-pipeline branch 2 times, most recently from 50b32c6 to 946978c Compare June 9, 2026 16:44
@Nirshal
Nirshal force-pushed the ho-release-gate-pipeline branch from e036171 to 69fb685 Compare June 19, 2026 15:44
@Nirshal
Nirshal force-pushed the ho-release-gate-pipeline branch from d3d3429 to dfc76bb Compare July 6, 2026 09:28
@Nirshal
Nirshal force-pushed the ho-release-gate-pipeline branch from a5ea2d1 to 14f5a1a Compare July 8, 2026 10:08
@Nirshal

Nirshal commented Jul 9, 2026

Copy link
Copy Markdown
Contributor Author

Addressing CodeRabbit nitpick: pin container image versions

All 7 occurrences of quay.io/konflux-ci/appstudio-utils:latest are now pinned using the tag+digest format (latest@sha256:32cda04...), following the Konflux recommended practice.

MintMaker (built on Renovate) will automatically open PRs to bump the digest when the :latest tag is updated in the registry. MintMaker scans all YAML files under .tekton/ weekly (Saturdays after 5 AM UTC), including custom pipelines like ours.

Nirshal added a commit to Nirshal/hypershift that referenced this pull request Jul 9, 2026
Pin quay.io/konflux-ci/appstudio-utils to tag+digest format as required
by Konflux policy. MintMaker (Renovate) will auto-bump the digest via
weekly PRs when the :latest tag is updated.

Add early validation for snapshot-name in extract-image to fail fast
before e2e execution.

Addresses CodeRabbit findings openshift#4 and openshift#6 on PR openshift#8602.

CNTRLPLANE-3434

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add a Konflux-based release gating pipeline that validates nightly
HyperShift Operator Snapshots before promoting them to downstream
managed services (ARO HCP first, ROSA and GCP to follow).

Pipeline flow:
- CronJob labels the latest Snapshot, Integration Service creates
  a PipelineRun via git-resolved ITS
- extract-image parses the Snapshot JSON for the HO container image
- run-e2e triggers blocking and informing Prow periodic jobs via
  Gangway, polls until completion (45 min initial delay, 10 min
  interval, 4h timeout)
- evaluate-results applies AND logic on blocking tests; informing
  failures are reported but do not block promotion
- create-release creates a Release CR on pass, triggering the
  managed pipeline to push the verified image to Quay

Stale promotion alerting (CNTRLPLANE-3451): on gate failure, both
finally tasks query KubeArchive for consecutive failure streaks and
send a dedicated Slack alert when the streak meets the configured
threshold.

Includes:
- Python stdlib-only modules: http_utils, prow_utils, slack_utils,
  kubearchive_utils, ho_release_gate
- Unit tests for all modules
- Mock utilities for end-to-end pipeline integration testing

JIRA: CNTRLPLANE-3434
OCPSTRAT: OCPSTRAT-3250

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@Nirshal
Nirshal force-pushed the ho-release-gate-pipeline branch from c2b1f20 to 31342d1 Compare July 9, 2026 15:55
@Nirshal
Nirshal marked this pull request as ready for review July 9, 2026 15:58
@openshift-ci openshift-ci Bot removed the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Jul 9, 2026
@openshift-ci
openshift-ci Bot requested review from enxebre and sjenning July 9, 2026 16:01
@hypershift-jira-solve-ci

Copy link
Copy Markdown
Contributor

Now I have all the data I need. Here is the final report:

Test Failure Analysis Complete

Job Information

Test Failure Analysis

Error

codecov/project: 42.66% (-0.94%) compared to 433a404 — FAILURE

@@            Coverage Diff             @@
##             main    #8602      +/-   ##
==========================================
- Coverage   43.60%   42.66%   -0.94%
==========================================
  Files         771      774       +3
  Lines       95806    97912    +2106
==========================================
  Hits        41778    41778
- Misses      51119    53225    +2106
  Partials     2909     2909

Summary

The codecov/project check failed because PR #8602 adds 13 new Python files under .tekton/lib/ containing ~2,106 coverable lines. Although the repository's codecov.yml has an ignore pattern ".*/**" intended to exclude dot-prefixed directories like .tekton/, Codecov is still counting 3 of these Python files (+2,106 lines) in the project coverage denominator. Since this is a Go project and no Python coverage reports are uploaded, all 2,106 lines register as uncovered (0 hits), dragging overall project coverage down from 43.60% to 42.66% — a 0.94% drop that triggers the default Codecov project gate (coverage must not decrease).

Root Cause

The root cause is a mismatch between the Codecov ignore pattern and Codecov's actual file-counting behavior for newly introduced Python files in a Go-only coverage project.

Step-by-step failure chain:

  1. PR adds Python files to a dot-directory: PR CNTRLPLANE-3434: add ho-release-gate pipeline for nightly promotion #8602 adds 13 Python files under .tekton/lib/ (source modules for a Tekton pipeline helper library: ho_release_gate.py, prow_utils.py, http_utils.py, slack_utils.py, kubearchive_utils.py, plus test and mock files).

  2. Ignore pattern should exclude them: The codecov.yml contains ".*/**" in the ignore: list, which in glob/fnmatch syntax matches any path starting with a dot-prefixed directory (.tekton/lib/anything). This pattern was designed to exclude dot-directories from coverage calculation.

  3. Pattern is not fully effective: Despite the ignore pattern, Codecov's file discovery still counts 3 new files contributing +2,106 coverable lines to the project denominator. This is evident from the coverage diff: Files: 771 → 774 (+3), Lines: 95,806 → 97,912 (+2,106).

  4. Zero coverage on those lines: Since the CI only uploads Go coverage reports (.out files from go test -coverprofile), Python files have no coverage data. All 2,106 newly counted lines are marked as misses: Misses: 51,119 → 53,225 (+2,106), Hits: 41,778 → 41,778 (unchanged).

  5. Project coverage gate fails: With no explicit coverage.project.default.threshold set in codecov.yml, Codecov uses its default auto target which fails when project coverage decreases. The 0.94% drop (43.60% → 42.66%) triggers the failure.

  6. Patch check passes (misleading): The codecov/patch check passes with "All modified and coverable lines are covered by tests" because Codecov's patch-level analysis treats these Python lines as outside the coverage domain (no Go coverage tool can cover them). The project-level check uses a different counting mechanism.

The fundamental issue is that Codecov is including Python source files discovered via repository scanning in the project coverage denominator, despite the ".*/**" ignore pattern that should exclude them. This is likely a Codecov glob-matching edge case where the pattern does not apply to files discovered through repo scanning in the same way it applies to files in the uploaded coverage report.

Recommendations
  1. Add an explicit .tekton/** ignore pattern to codecov.yml — this is the most direct fix and avoids relying on the ".*/**" glob matching behavior:

    ignore:
      - ".*/**"
      - ".tekton/**"        # ← add this explicit pattern
      - "**/*.py"           # ← or exclude all Python files entirely
  2. Add a coverage threshold to tolerate minor drops — configure the project target to allow small decreases, preventing false failures from non-Go files:

    coverage:
      project:
        default:
          threshold: 1%     # Allow up to 1% decrease
  3. Add "**/*.py" to the ignore list — since this is a Go project with no Python coverage instrumentation, all .py files should be excluded from coverage calculation regardless of location:

    ignore:
      - "**/*.py"
  4. Re-trigger CI on the latest commit — the Codecov report is stale (computed against c2b1f20 while the PR head is now 31342d1). After fixing codecov.yml, push or re-trigger to get a fresh report.

Option 3 ("**/*.py") is the recommended approach — it's the broadest fix that prevents any future Python file additions (in .tekton/ or elsewhere) from affecting Go coverage metrics.

Evidence
Evidence Detail
Check Run ID 86165231728 — concluded failure at 2026-07-09T15:56:39Z
Coverage Drop 43.60% → 42.66% (−0.94%) against base 433a404
New Files Counted +3 files, +2,106 coverable lines, +0 hits, +2,106 misses
Files Changed 16 files: 13 .py, 1 .md, 2 .yaml — all under .tekton/
Ignore Pattern ".*/**" in codecov.yml — should match .tekton/ but is not fully effective
No Explicit Threshold codecov.yml has no coverage.project section; default auto fails on any decrease
Patch Check codecov/patch passed — "All modified and coverable lines are covered by tests"
Stale Report Codecov warns: "Current head c2b1f20 differs from pull request most recent head 31342d1"
Flag Coverage Drops cmd-support: −0.80%, hypershift-operator: −2.89% (denominator inflation)
Python Line Breakdown 949 source lines + 1,454 test/mock lines = 2,403 raw lines (~2,106 coverable)

@Nirshal

Nirshal commented Jul 9, 2026

Copy link
Copy Markdown
Contributor Author

/area ci-tooling

The clone-lib task and PipelineRun git resolver were pointing to
the fork (Nirshal/hypershift) instead of the upstream repo
(openshift/hypershift).

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
Commit-Message-Assisted-by: Claude (via Claude Code)

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 6

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.tekton/lib/ho_release_gate.py:
- Around line 630-701: `check_and_build_stale_payload` can crash the finally
notification path if `datetime.fromisoformat` receives an invalid
`oldest_created` timestamp. Add error handling around the timestamp parsing in
this helper so a bad KubeArchive value is treated like a skipped stale check,
with a clear log message, and keep the existing `build_stale_notification` flow
unchanged for valid timestamps.

In @.tekton/lib/kubearchive_utils.py:
- Around line 16-84: fetch_pipelineruns currently only processes the first
KubeArchive list response, so it can miss older PipelineRuns when the API
paginates results. Update the fetch loop in fetch_pipelineruns to keep
requesting additional pages using metadata.continue (or an explicit limit plus
continue token) until no continue token is returned, then merge all items before
sorting and returning the runs. Keep the existing parsing and logging behavior
in fetch_pipelineruns, but make sure every page’s items are included so
streak_days and stale checks see the full history.

In @.tekton/lib/mock/test_util_mock.py:
- Around line 256-266: The docstring in the stale alert test helper is
inconsistent with the generation logic: the oldest synthetic PipelineRun is
actually `threshold_days + 1` days ago, not `threshold_days` ago. Update the
documentation in `test_util_mock.py` around the stale streak generator (and the
related copy in the other referenced block) to match the behavior of `num_runs`,
`days_ago`, and the stale promotion helper so engineers debugging the alert see
accurate timing semantics.

In @.tekton/lib/prow_utils.py:
- Around line 160-167: `get_prow_job_status()` currently treats `status == 0` as
a generic error, which makes `poll_until_complete()` fail fast on transient
transport issues. Update the status classification logic in
`get_prow_job_status()` so `0` is handled as retryable (either alongside the
`"rate_limited"` path or as a distinct connection-failure result), and ensure
the caller path in `poll_until_complete()` continues polling instead of marking
the job terminal for this case.
- Around line 58-69: Stop retrying 5xx responses for the trigger POST in
trigger_prow_job, since a lost response can duplicate Prow executions. Update
should_retry in prow_utils.py to only retry transport failures and 429 for this
path, or otherwise make the Gangway POST idempotent before allowing 5xx retries.
Use the existing http_request_with_retry call and its should_retry helper as the
place to change this behavior.

In @.tekton/lib/tests/test_prow_utils.py:
- Around line 95-106: The test_passes_correct_payload case in test_prow_utils
should actually verify the POST body built by trigger_prow_job, since it
currently only checks retries/backoff and never asserts the payload. Simplify
the mock call inspection around http_request_with_retry, capture the body passed
by trigger_prow_job, and add assertions that job_name and env_overrides are
present in that payload with the expected values.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 603cb9ef-4cac-41e9-bd66-91288cc84b52

📥 Commits

Reviewing files that changed from the base of the PR and between 58a3243 and f75da4b.

📒 Files selected for processing (16)
  • .tekton/lib/README.md
  • .tekton/lib/ho_release_gate.py
  • .tekton/lib/http_utils.py
  • .tekton/lib/kubearchive_utils.py
  • .tekton/lib/mock/__init__.py
  • .tekton/lib/mock/test_util_mock.py
  • .tekton/lib/prow_utils.py
  • .tekton/lib/slack_utils.py
  • .tekton/lib/tests/__init__.py
  • .tekton/lib/tests/test_ho_release_gate.py
  • .tekton/lib/tests/test_http_utils.py
  • .tekton/lib/tests/test_kubearchive_utils.py
  • .tekton/lib/tests/test_prow_utils.py
  • .tekton/lib/tests/test_slack_utils.py
  • .tekton/pipelines/ho-release-gate-run.yaml
  • .tekton/pipelines/ho-release-gate.yaml
🚧 Files skipped from review as they are similar to previous changes (3)
  • .tekton/pipelines/ho-release-gate-run.yaml
  • .tekton/lib/tests/test_slack_utils.py
  • .tekton/pipelines/ho-release-gate.yaml

Comment thread .tekton/lib/ho_release_gate.py
Comment thread .tekton/lib/kubearchive_utils.py Outdated
Comment thread .tekton/lib/mock/test_util_mock.py Outdated
Comment thread .tekton/lib/prow_utils.py
Comment thread .tekton/lib/prow_utils.py Outdated
Comment thread .tekton/lib/tests/test_prow_utils.py
- Guard datetime.fromisoformat against invalid timestamps in stale
  check to prevent finally task crash (ho_release_gate.py)
- Add time filter (creationTimestampAfter) and explicit limit to
  KubeArchive query to ensure all recent runs are fetched regardless
  of server-side sort order (kubearchive_utils.py)
- Treat connection failures (status 0) as retryable in
  get_prow_job_status so polling continues on transient errors
  (prow_utils.py)
- Fix test_passes_correct_payload to verify POST body content
  (test_prow_utils.py)
- Fix docstring timing semantics in stale mock (test_util_mock.py)

JIRA: CNTRLPLANE-3434

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What are these changes?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mock/ directory contains test_util_mock.py, a utility module that generates synthetic test data for end-to-end pipeline integration testing (fake Snapshots, simulated Gangway/Prow responses, stale streak generators). The __init__.py makes it a Python package so the utilities can be imported. These are not unit tests (those are in tests/) but helpers for validating the full pipeline flow. I left them in the PR because I found them very useful to test some corner cases, and I thought they might have value for future maintenance. That said, working with these requires temporarily altering the Tekton pipeline, which is something possible only on an open PR (which still requires the ITS to point to the PR branch instead of main), so I am open to removing these and archiving them locally on my computer, just in case I need them in the future. Waiting for feedback from you on this.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right — the mock package is a self-contained toolkit for manual integration testing, not a unit test suite. fetch_pipelineruns_mock_stale_long is documented with usage instructions in the module docstring and is part of the mock API surface. "Not called from tests/" doesn't make it dead code here. Withdrawing.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right — the mock package is a self-contained toolkit for manual integration testing, not a unit test suite. fetch_pipelineruns_mock_stale_long is documented with usage instructions in the module docstring and is part of the mock API surface. "Not called from tests/" doesn't make it dead code here. Withdrawing.

@bryan-cox bryan-cox left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: feat(ci): add ho-release-gate pipeline for nightly promotion

Well-structured Tekton pipeline with solid Python library design. Two blocking issues (both the same class: feature branch references that need updating to main), several hardening suggestions, and a few questions. See inline comments.

Category Count
Blocking 2
Suggestions 6
Nits 1
Questions 2
Praise 3

Praise:

  • Excellent defensive design with default-first result initialization — prevents silent failures on OOMKill
  • Comprehensive test suite with well-structured edge cases
  • stdlib-only constraint is well-justified with clear upgrade path documented in README

Once the two branch references are corrected, this should be in good shape for a second pass.

- name: url
value: https://github.com/openshift/hypershift
- name: revision
value: ho-release-gate-pipeline

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocking] This references the feature branch ho-release-gate-pipeline. Once merged and the branch is deleted, the git resolver will fail to find the Pipeline definition.

Suggested change
value: ho-release-gate-pipeline
value: main

If the ITS configuration overrides this value at runtime, please add a comment here stating that.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6028dd6. Changed to openshift/hypershift + main.

Comment thread .tekton/pipelines/ho-release-gate.yaml Outdated
echo "=== clone-lib ==="
echo "Running as: $(oc whoami)"
REPO_URL="https://github.com/openshift/hypershift.git"
BRANCH="ho-release-gate-pipeline"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocking] Same issue as the PipelineRun template — this sparse checkout of .tekton/lib/ will break once the feature branch is deleted post-merge.

Suggested change
BRANCH="ho-release-gate-pipeline"
BRANCH="main"

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6028dd6. Changed to openshift/hypershift + main.

Comment thread .tekton/pipelines/ho-release-gate.yaml Outdated
with open("$(results.results-json.path)", "w") as f:
f.write("[]")

gangway_url = "https://gangway-ci.apps.ci.l2s4.p1.openshiftapps.com/v1/executions"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] This Gangway URL is hardcoded. The CI cluster has migrated before (app.ci → build0x). Consider making this a pipeline parameter with this as the default, so a cluster move doesn't require a code change.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. The Gangway URL is hardcoded in a single place (the run-e2e inline script). I will extract it as a pipeline parameter with the current value as default, so a cluster migration only requires updating the ITS definition.


env_overrides = {
"OVERRIDE_IMAGE_HYPERSHIFT_OPERATOR": ho_image,
"OVERRIDE_IMAGE_HYPERSHIFT_TESTS": "registry.ci.openshift.org/hypershift/hypershift-tests:latest"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] Using :latest for the test image means the gate runs whatever test binary happens to be latest, which may not correspond to the HO image being validated. If the Snapshot contains a test image component, it would be more correct to extract and use that. If :latest is intentional (e.g., tests are always forward-compatible), a comment explaining why would help future readers.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a known limitation that I already flagged to the team (see Slack thread). The Konflux Snapshot only contains the hypershift-operator component. The test image (hypershift-tests) is built by the OpenShift CI pipeline, not by Konflux, so it is not part of the Snapshot and there is no straightforward way to extract the matching version from another source. Using :latest is intentional given this constraint. I will add a comment in the code explaining why.

pipeline_run = "$(context.pipelineRun.name)"
results_json = os.environ["RESULTS_JSON"]
webhook_url = os.environ["SLACK_WEBHOOK_URL"]
konflux_base = "https://konflux-ui.apps.stone-prd-rh01.pg1f.p1.openshiftapps.com/ns/crt-redhat-acm-tenant/applications/hypershift-operator"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] Since ARO HCP is the pilot and ROSA/GCP HCP follow using the same template, consider extracting the namespace (crt-redhat-acm-tenant) and Konflux app URL as pipeline parameters. This avoids forking the pipeline when onboarding new platforms — their ITS definitions can supply different values.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The namespace crt-redhat-acm-tenant is the Konflux tenant where all resources live (Snapshots, ReleasePlans, Release CRs) -- it is not specific to ARO HCP. When onboarding ROSA or GCP, each will have its own ITS definition and ReleasePlan, but they will all live in the same tenant namespace. The pipeline is already designed to be reusable across managed services through its parameters (gate-label, release-plan-name, e2e-blocking-job-names, etc.), so no forking is needed. The Konflux app URL is only used for building links in Slack notifications and is also tenant-level, not service-specific.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Makes sense — the namespace and URL are tenant-level constants shared across all managed services, not service-specific values. The pipeline is already parameterized at the right level (gate-label, release-plan-name, e2e-blocking-job-names). Extracting these would add complexity for a scenario that won't happen. Withdrawing.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Makes sense — the namespace and URL are tenant-level constants shared across all managed services, not service-specific values. The pipeline is already parameterized at the right level (gate-label, release-plan-name, e2e-blocking-job-names). Withdrawing.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

test reply

release_name, snapshot, pipeline_run, konflux_base)

success = send_slack_message(webhook_url, payload)
if not success:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] When the Slack notification fails, this task still exits 0, making the failure invisible in PipelineRun status. Consider sys.exit(1) so the PipelineRun shows a partial failure, or write a result indicating notification status. I understand the rationale of not blocking pipeline completion from a finally task, but silent notification failure means operators won't know alerts are broken.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I intentionally chose to exit 0 on notification failure. If notify-slack exits non-zero, the PipelineRun is marked as Failed in Tekton and KubeArchive. When an operator notices that notifications are missing and starts investigating, they would see all PipelineRuns marked as Failed and would have to open each one to determine whether the failure was a real gate failure or just a broken notification. By keeping the exit code tied to the gate verdict only, the PipelineRun status directly answers "did the release succeed?" without ambiguity. At that point the operator already knows notifications are broken and can check the task logs of the latest run to understand why.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good rationale. PipelineRun status = gate verdict is the right invariant for a finally task — mixing in notification health would make triage harder, not easier. Withdrawing.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good rationale. PipelineRun status = gate verdict is the right invariant for a finally task — mixing in notification health would make triage harder, not easier. Withdrawing.

pipeline_run, konflux_base, gate_label)

success = send_slack_message(webhook_url, payload)
if not success:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[question] The PR description says notify-slack and notify-slack-error are "mutually exclusive." Is there a scenario where neither fires? A short comment in the YAML explaining the mutual exclusivity logic (which task fires under which condition) would help future maintainers.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that this is a bit counter-intuitive. The notify-slack-error is meant to be a "catch-all" before the release task is reached, so every error before reaching that point will be reported by that. The reason why we also have error handling in the other finally task is that after create-release, more information about the errors becomes available, so the message can be more helpful. The mutual exclusivity exists and relies on Tekton's task status propagation: (1) If the pipeline reaches evaluate-results, then create-release runs (Succeeded or Failed), notify-slack fires (it receives all resolved params), and notify-slack-error is skipped (create-release.status != None). (2) If the pipeline fails before evaluate-results (e.g. extract-image or run-e2e crash), create-release is skipped (status=None), its result params are unresolved so notify-slack cannot run, and notify-slack-error fires instead. There is no scenario where neither fires. I will add a YAML comment block above the finally section explaining this logic.


The output format is a JSON array of objects with keys:
job (full name), result, url, type. Uses compact separators
to minimize size (Tekton results have a 4KB limit).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] With multiple blocking + informing jobs, each carrying a full Prow URL (~150 chars), the result could exceed 4KB. Consider either:

  1. Adding a size check with truncation fallback
  2. Writing results to a workspace file instead of a Tekton result
  3. Truncating URLs in the JSON (full URLs are already printed to stdout)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, this is a real risk. I will switch to your option 2: write the results JSON to a workspace file instead of a Tekton result. The shared workspace is already mounted on all tasks (we use it for the Python libraries), so both evaluate-results and notify-slack can read from it directly. I will also remove the Tekton result declaration to avoid confusion.

Comment thread .tekton/lib/kubearchive_utils.py Outdated

__all__ = ["fetch_pipelineruns", "build_pipelinerun_url"]

KUBEARCHIVE_API_BASE = (

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] This couples the library to a specific production cluster. Consider making it configurable via env var for staging/alternative environments:

KUBEARCHIVE_API_BASE = os.environ.get(
    "KUBEARCHIVE_API_BASE",
    "https://kubearchive-api-server-product-kubearchive"
    ".apps.stone-prd-rh01.pg1f.p1.openshiftapps.com"
)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. I will make it configurable via env var with the current value as default, as you suggested.

Comment thread .tekton/pipelines/ho-release-gate.yaml Outdated
threshold = int(os.environ["STALE_THRESHOLD_DAYS"])
stale = check_and_build_stale_payload(
token, its_scenario, threshold, pipeline_run,
konflux_base, "crt-redhat-acm-tenant", gate_label)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[question] Since generateName is used with oc create in the create-release task, re-triggering the pipeline for the same Snapshot creates a duplicate Release CR. Is there deduplication upstream (ReleasePlanAdmission, managed pipeline), or should there be a check here?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no deduplication at the Release CR creation level, so a duplicate CR will be created. However, the managed release pipeline is designed to be idempotent. Early in execution, the filter-already-released-advisory-images task checks if the container image digests from the Snapshot have already been pushed to the target repository. If so, it sets a skip_release flag that causes Tekton to skip all downstream tasks (push-snapshot, rh-sign-image, create-pyxis-image, create-advisory, etc.). Individual tasks also check their own domains separately as an additional safety net. In practice, a duplicate Release CR will start the pipeline but it will short-circuit after the initial setup tasks without altering the target registry or creating duplicate advisories. If you want, we could add a pre-creation check, but this will add complexity and could also not be straightforward, because we could have several Release CRs for the same Snapshot when we add more managed services (like ROSA) and in that case it would be further difficult and prone to errors to correctly check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misread the failure mode — generateName + oc create produces a new CR each time, so there's no collision. And the managed release pipeline's filter-already-released-advisory-images task handles idempotency downstream. No dedup needed here. Withdrawing.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misread the failure mode — generateName + oc create produces a new CR each time, so there's no collision. The managed release pipeline's filter-already-released-advisory-images handles idempotency downstream. Withdrawing.

Nirshal and others added 5 commits July 16, 2026 14:34
Switch from OVERRIDE_IMAGE_HYPERSHIFT_OPERATOR (ci-operator ImageStream
mechanism, broken by race condition) to the MULTISTAGE_PARAM_OVERRIDE_
transport variable that passes the image directly to the hypershift-install
step parameter.

Requires openshift/release#81877 to be merged first.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Point the PipelineRun git resolver and clone-lib back to
Nirshal/hypershift so the pipeline can run against the feature
branch and validate the OVERRIDE_HYPERSHIFT_OPERATOR_IMAGE flow
end-to-end. Will be switched to openshift/hypershift + main after
validation.

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
The PipelineRun git resolver and clone-lib sparse checkout were
referencing the fork branch (Nirshal/hypershift, ho-release-gate-pipeline).
After merge, the feature branch will be deleted and both would fail.

Switch to openshift/hypershift + main so the pipeline resolves from
upstream permanently.

Signed-off-by: Alessandro Rossi <alesross@redhat.com>
…ve docs

- Extract Gangway URL as pipeline parameter with default value
- Extract KubeArchive URL as pipeline parameter, pass through
  function arguments instead of module-level env var read
- Add comment explaining intentional :latest for test image
- Improve finally tasks comment explaining Tekton skip mechanism
- Update tests for new function signatures

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Tekton task results have a 4KB limit. With many blocking and informing
jobs, each carrying a full Prow URL, the results JSON can exceed that
limit and crash the task at runtime. Write results to a file on the
shared workspace instead, and have downstream tasks read from it.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@bryan-cox bryan-cox left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Second-round review. 6 inline comments — 4 carry forward from the first round (still open), 2 are new findings.

with open("$(workspaces.shared.path)/results.json") as f:
results_json = f.read()
webhook_url = os.environ["SLACK_WEBHOOK_URL"]
konflux_base = "https://konflux-ui.apps.stone-prd-rh01.pg1f.p1.openshiftapps.com/ns/crt-redhat-acm-tenant/applications/hypershift-operator"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still open from round 1 — hardcoded Konflux URL + namespace

konflux_base and the namespace crt-redhat-acm-tenant are baked into the pipeline YAML. If the tenant or Konflux instance changes, every pipeline that inlines this needs an update.

Consider lifting both into pipeline parameters (with the current values as defaults) so they can be overridden at the PipelineRun level without touching the Pipeline definition.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Alessandro's response is correct — the namespace and URL are tenant-level constants shared across all managed services, not service-specific. The pipeline is already parameterized at the right level. Withdrawing.


success = send_slack_message(webhook_url, payload)
if not success:
print("ERROR: Slack notification failed after all retries", flush=True)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still open from round 1 — silent exit 0 on Slack failure

When Slack notification fails, the task prints the error but still exits 0. Downstream consumers (humans, dashboards) have no way to tell that notification was lost.

Consider writing a task result (e.g. $(results.slack-status.path)) with "failed" so callers can detect it, even if the task itself should not block the pipeline.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Alessandro's rationale holds. PipelineRun status = gate verdict is the right invariant for a finally task — mixing in notification health would make triage harder, not easier. Withdrawing

print(f"Release: {release_name}", flush=True)

if not gate_passed:
its_scenario = os.environ["ITS_SCENARIO"]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New — duplicated stale-check block

The ~15-line stale-check block (its_scenario = … through send_slack_message(stale_payload, …)) is copy-pasted between notify-slack and notify-slack-error. This is the same logic with the same env vars — a classic Duplicated Code smell.

Extract it into a shared shell function or a separate Task so changes only need to happen in one place.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The duplicated block is the parameter preparation for check_and_build_stale_payload(), which is where the actual shared logic lives (already extracted in ho_release_gate.py). The env var reads and token acquisition are the glue that feeds values into that function.

Extracting this glue into a helper would create a script-like function that reads from os.environ and calls subprocess internally. This would be difficult to test in isolation (it's not a pure function, it's effectively a script), and it would break the design pattern we follow throughout the pipeline: testable, reusable functions in the Python library, called by thin inline scripts in the YAML. Adding a helper in between would introduce a layer of indirection (scripts calling scripts calling functions) without reducing complexity.

Since the two finally tasks run as separate Tekton pods with no shared execution context, each must independently prepare its own parameters. I think the current factoring is the right one, but happy to chat about it if you see it differently.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Misread the failure mode — generateName + oc create produces a new CR each time, no collision. The managed release pipeline's filter-already-released-advisory-images handles idempotency downstream. Withdrawing.

"OVERRIDE_IMAGE_HYPERSHIFT_TESTS": "registry.ci.openshift.org/hypershift/hypershift-tests:latest"
}

jobs = trigger_all_jobs(blocking, informing, gangway_url, token, env_overrides)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New — run-e2e can exit non-zero despite spec "always exits 0"

trigger_all_jobs() raises ValueError on missing env vars, and this script: block has no try/except around it. If GANGWAY_URL or a job-list var is unset, the task exits non-zero and the pipeline fails hard — violating the spec requirement that run-e2e always exits 0, deferring pass/fail to evaluate-results.

Wrap the call in a try/except that catches ValueError, logs it, and writes a sentinel to $(results.*) so evaluate-results can report the real cause.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The pipeline has two distinct error categories by design, each with its own reporting path:

  1. Test failures (normal operation): one or more Prow jobs fail. run-e2e exits 0, writes structured results to the workspace, evaluate-results determines the gate verdict, and notify-slack sends a detailed notification with per-job status, URLs, and stale promotion alerts. This is the path optimized for actionable detail.

  2. Infrastructure errors (exceptional): a task crashes due to misconfiguration, transient cluster issues, API outages, etc. notify-slack-error fires and alerts the team that the pipeline itself broke. Troubleshooting goes through PipelineRun logs, which are the only place that can provide full context for unexpected failures.

These two paths are intentionally separate. Adding error handling inside run-e2e to catch infrastructure errors and route them through notify-slack would blur this boundary: we would have two error reporting paths for infrastructure failures (one partial in run-e2e, one generic in notify-slack-error), and the partial one could never be exhaustive because it only covers one task out of four. A crash in clone-lib, extract-image, or evaluate-results would still go through notify-slack-error regardless. The result would be added complexity without added coverage.

On the specific scenario (missing env vars): GANGWAY_URL, HO_IMAGE, BLOCKING_JSON, and INFORMING_JSON are all injected via explicit Tekton param bindings in the task definition, and GANGWAY_TOKEN comes from a secretKeyRef. If any of these are missing, Tekton fails the pod scheduling before the script runs - the os.environ reads are never reached.

echo "ReleasePlan: ${RELEASE_PLAN}"

set +e
RELEASE_NAME=$(oc create -f - -o jsonpath='{.metadata.name}' <<EOFRELEASE

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still open from round 1 — duplicate Release CR on re-trigger

If the pipeline is re-triggered for the same snapshot, oc create -f - will attempt to create a Release CR that may already exist, causing a failure. Is there deduplication logic elsewhere (e.g. a unique name derived from the snapshot SHA), or should this use oc apply / create-if-not-exists?

return runs


def fetch_pipelineruns_mock_stale_long(token, namespace, label_selector):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: fetch_pipelineruns_mock_stale_long is exported in __all__ but never referenced in any test file. Dead code — either add test coverage that uses it or remove it.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mock/ package is a self-contained toolkit for manual integration testing, not a unit test suite. The function is documented with usage instructions in the module docstring. "Not called from tests/" doesn't make it dead code. Withdrawing.

@bryan-cox

Copy link
Copy Markdown
Member

Replying to Alessandro's four responses from the first round — he's right on all of them, withdrawing those items.

Mock dead code (fetch_pipelineruns_mock_stale_long): The mock/ package is a self-contained toolkit for manual integration testing, not a unit test suite. The function is documented with usage instructions in the module docstring and is part of the mock API surface. "Not called from tests/" doesn't make it dead code here. Withdrawing.

Konflux URL + namespace: The namespace and URL are tenant-level constants shared across all managed services, not service-specific values. The pipeline is already parameterized at the right level (gate-label, release-plan-name, e2e-blocking-job-names). Withdrawing.

Silent exit 0 on Slack failure: PipelineRun status = gate verdict is the right invariant for a finally task — mixing in notification health would make triage harder, not easier. Withdrawing.

Duplicate Release CR: I misread the failure mode — generateName + oc create produces a new CR each time, so there's no collision. The managed release pipeline's filter-already-released-advisory-images handles idempotency downstream. Withdrawing.

The two new findings from round 2 (duplicated stale-check block and run-e2e ValueError path) still stand as independent items. Sorry for the noise on the re-opened ones.

@openshift-ci openshift-ci Bot added the lgtm Indicates that a PR is ready to be merged. label Jul 21, 2026
@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Pipeline controller notification

No second-stage tests were triggered for this PR.

This can happen when:

  • The changed files don't match any pipeline_run_if_changed patterns
  • All files match pipeline_skip_if_only_changed patterns
  • No pipeline-controlled jobs are defined for the main branch

Use /test ? to see all available tests.

@openshift-ci

openshift-ci Bot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: bryan-cox, Nirshal

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@openshift-ci openshift-ci Bot added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Jul 21, 2026
@Nirshal Nirshal changed the title feat(ci): add ho-release-gate pipeline for nightly promotion CNTRLPLANE-3434: add ho-release-gate pipeline for nightly promotion Jul 21, 2026
@openshift-ci-robot openshift-ci-robot added the jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. label Jul 21, 2026
@openshift-ci-robot

openshift-ci-robot commented Jul 21, 2026 •

Copy link
Copy Markdown

@Nirshal: This pull request references CNTRLPLANE-3434 which is a valid jira issue.

Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the epic to target the "5.0.0" version, but no target version was set.

Details

In response to this:

Summary

Adds the Tekton pipeline and PipelineRun template for the HyperShift Operator nightly release gating pipeline, triggered via IntegrationTestScenario (ITS).

Pipeline flow

CronJob (3:15 UTC nightly)
 |-- Resolves latest Konflux Snapshot (auto-released, push build)
 |-- Labels Snapshot with test.appstudio.openshift.io/run=<its-name>

Integration Service
 |-- Detects labeled Snapshot, creates PipelineRun from template
 |-- Injects SNAPSHOT JSON payload + ITS params into PipelineRun

PipelineRun (ho-release-gate):
 extract-image     - Parses SNAPSHOT JSON to extract HO image and commit SHA.
                     Gets snapshot name from PipelineRun label.
 run-e2e           - Triggers Prow periodic e2e jobs via gangway API.
                     Two categories: blocking (gate fails if any fail) and
                     informing (reported but do not block the gate).
                     45min initial delay, then polls every 10 min with 30s stagger.
                     Overrides: MULTISTAGE_PARAM_OVERRIDE_OVERRIDE_HYPERSHIFT_OPERATOR_IMAGE
                                (bypasses ci-operator ImageStream race condition),
                                OVERRIDE_IMAGE_HYPERSHIFT_TESTS.
                     Always exits 0 - writes per-job results JSON (with type field).
 evaluate-results  - Reads results, applies AND logic on blocking tests only.
                     Informing failures logged as warnings.
                     Always exits 0 - writes gate-passed result (true/false).
 create-release    - If gate passed: creates Release CR referencing the ReleasePlan
                     and Snapshot, writes release name to result.
                     If gate failed: writes "N/A" to result, exits 1
                     (marks pipeline as Failed in UI).
 finally:
   notify-slack       - Fires on every run (pass or fail).
                        Three visual states:
                          Green (#2E7D32): gate passed, Release created.
                          Orange (#F57C00): gate passed, Release creation failed.
                          Red (#D32F2F): gate failed.
                        Gate-label in header, grouped results (blocking + informing).
                        On gate failure: checks for stale promotion streak via
                        KubeArchive. If streak >= threshold, sends stale alert
                        instead of normal failure notification.
   notify-slack-error - Fallback for catastrophic failures (DAG failure).
                        Zero result dependencies, fires when create-release
                        did not run. Gate-label in header.
                        Also checks for stale promotion streak via KubeArchive.
                        Sends stale alert or generic error notification.

Release (on gate pass):
 create-release task creates Release CR explicitly.
 ReleasePlan targets rhtap-releng-tenant (managed workspace).
 Managed pipeline pushes image to verified Quay repo.

Stale promotion alerting (CNTRLPLANE-3451)

When the gate fails, both notify-slack and notify-slack-error query the
KubeArchive REST API for historical PipelineRuns matching the ITS scenario label.
The consecutive failure streak is measured in days (difference between now and the
oldest consecutive failed run, inclusive of today).

If streak_days >= stale-threshold-days, the normal notification is replaced with
a stale promotion alert containing:

  • Header with gate label and streak duration
  • Fields: Gate, Threshold, Streak, Current PipelineRun
  • History of up to 10 most recent failed runs with clickable PipelineRun links,
    dates, and failure reasons (Failed, PipelineRunTimeout, CouldntGetPipeline, etc.)
  • Truncation summary for streaks longer than 10 runs
  • Footer with investigation guidance

The stale check is safe to skip: if KubeArchive is unreachable or returns no data,
the normal notification is sent instead.

Notification reliability

All tasks use default-first initialization: each task writes safe default values
to its result paths at the start of its script, before any logic. If a task starts
but crashes mid-execution (e.g., OOMKilled), results are still initialized.

Two mutually exclusive finally tasks ensure a notification is always sent:

  • notify-slack: detailed notification, depends on task results (skipped if results uninitialized)
  • notify-slack-error: generic fallback, zero result dependencies, fires only when create-release was skipped by DAG failure

Both finally tasks perform the stale promotion check independently, so a stale alert
is sent regardless of which notification path fires.

Files

File Description
.tekton/pipelines/ho-release-gate.yaml Pipeline definition
.tekton/pipelines/ho-release-gate-run.yaml PipelineRun template (referenced by ITS via git resolver)
.tekton/lib/ho_release_gate.py Pipeline orchestration: image extraction, job triggering/polling, gate evaluation, notifications, stale alerting
.tekton/lib/kubearchive_utils.py KubeArchive REST API: fetch historical PipelineRuns
.tekton/lib/prow_utils.py Gangway API: trigger jobs, resolve URLs, poll status
.tekton/lib/slack_utils.py Slack webhook: send messages, build Block Kit payloads
.tekton/lib/http_utils.py Low-level HTTP wrapper with retry logic (stdlib only)
.tekton/lib/README.md Library documentation: architecture, constraints, runtime environment
.tekton/lib/tests/ Unit tests for all modules (unittest + unittest.mock, stdlib only)
.tekton/lib/mock/ Mock utilities for pipeline integration testing

Pipeline parameters

Parameter Source Description
SNAPSHOT Integration Service (auto-injected) JSON payload of the Konflux Snapshot
e2e-blocking-job-names ITS params JSON array of blocking Prow periodic job names
e2e-informing-job-names ITS params JSON array of informing Prow periodic job names
gate-label ITS params Label identifying which service gate (e.g. ARO HCP)
release-plan-name ITS params Name of the ReleasePlan to reference when creating the Release CR
stale-threshold-days ITS params Number of consecutive failure days before triggering stale alert (default: 3)

Resources

All resources live in namespace crt-redhat-acm-tenant unless noted.

Resource Name Managed via
Pipeline + PipelineRun template .tekton/pipelines/ho-release-gate*.yaml This PR
ReleasePlanAdmission redhat-hypershift-operator-ho-release-gate-aro-hcp (in rhtap-releng-tenant) GitOps (!18934 - Merged)
ReleasePlan hypershift-operator-ho-release-gate-aro-hcp GitOps (!18938 - Merged)
CronJob hypershift-operator-nightly-promotion GitOps (!18938 - Merged)
IntegrationTestScenario hypershift-ho-release-gate-aro-hcp GitOps (!19261 - Merged)
ServiceAccount nightly-promotion-sa GitOps (!18934 - Merged)
RBAC (snapshot labeling) nightly-promotion-sa -> konflux-tester-internalbot-actions GitOps (!19870 - Merged)
RBAC (release creation) nightly-promotion-sa -> konflux-releaser-bot-actions via taskRunSpecs GitOps (!19870 - Merged)
Secret gangway-token (Prow auth) Manual (not in GitOps)
Secret slack-webhook (Slack notifications) Manual (not in GitOps)

Release mechanism

The pipeline uses explicit Release CR creation (not auto-release):

  1. The create-release task creates a Release object referencing the ReleasePlan (name passed via ITS param release-plan-name) and the tested Snapshot
  2. The ReleasePlan targets rhtap-releng-tenant (managed workspace)
  3. The RPA (redhat-hypershift-operator-ho-release-gate-aro-hcp) picks up the Release
  4. The managed pipeline (rh-push-to-external-registry) pushes the validated image to the verified Quay repo with platform-prefixed tags (aro-hcp-latest, aro-hcp-latest-{{ timestamp }}, aro-hcp-{{ git_sha }}, etc.)

Target repo: quay.io/redhat-services-prod/crt-redhat-acm-tenant/hypershift/hypershift-operator-verified

Note: Auto-release was disabled (!19296 - Merged) because with contexts: disabled on the ITS, Integration Service was auto-releasing every normal push snapshot through our ReleasePlan, bypassing the release gate entirely.

Remaining work

Item Status Notes
HO image override fix Blocked Depends on openshift/release#81877 (adds OVERRIDE_HYPERSHIFT_OPERATOR_IMAGE step parameter + Gangway transport variable to hypershift-install)
Pipeline source migration On merge ITS resolverRef will point to openshift/hypershift:main once this PR merges (!19913 - Draft)
Regression analysis Future Component-readiness tracking (deads2k requirement)

Related

Summary by CodeRabbit

  • New Features
  • Added a new Tekton release gate pipeline that parses the provided snapshot, runs configured Gangway Prow jobs, and computes a gate verdict.
  • Release creation is skipped when blocking checks don’t pass.
  • Added Slack notifications for normal results, errors, and stale-promotion conditions; notifications run even on failures.
  • Documentation
  • Documented the shared stdlib-only release-gate helper library and how to run its tests.
  • Tests
  • Added unit test coverage for HTTP retries, Prow/Gangway helpers, gate evaluation, Slack payloads, and stale detection logic.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@Nirshal

Nirshal commented Jul 21, 2026

Copy link
Copy Markdown
Contributor Author

/verified later @Nirshal

@openshift-ci-robot openshift-ci-robot added verified-later verified Signifies that the PR passed pre-merge verification criteria labels Jul 21, 2026
@openshift-ci-robot

Copy link
Copy Markdown

@Nirshal: This PR has been marked to be verified later by @Nirshal.

Details

In response to this:

/verified later @Nirshal

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@Nirshal

Nirshal commented Jul 21, 2026

Copy link
Copy Markdown
Contributor Author

/retest ci/prow/okd-scos-images

@Nirshal

Nirshal commented Jul 21, 2026

Copy link
Copy Markdown
Contributor Author

/test okd-scos-images

@openshift-ci

openshift-ci Bot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

@Nirshal: all tests passed!

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@openshift-merge-bot
openshift-merge-bot Bot merged commit 65c836d into openshift:main Jul 21, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. area/ci-tooling Indicates the PR includes changes for CI or tooling jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. lgtm Indicates that a PR is ready to be merged. verified Signifies that the PR passed pre-merge verification criteria verified-later

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants