Repository navigation
Conversation
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.23-periodics-e2e-aws-private-sg-cleanup |
|
Skipping CI for Draft Pull Request. |
|
@zhfeng: This pull request references Jira Issue OCPBUGS-74960, which is valid. 3 validation(s) were run on this bug
Requesting review from QA contact: The bug has been updated to refer to the pull request using the external bug tracker. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
@zhfeng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
@openshift-ci-robot: GitHub didn't allow me to request PR reviews from the following users: zhfeng. Note that only openshift members and repo collaborators can review this PR, and authors cannot review their own PRs. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
[APPROVALNOTIFIER] This PR is NOT APPROVED This pull-request has been approved by: zhfeng The full list of commands accepted by this bot can be found here. DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.23-periodics-e2e-aws-private-sg-cleanup |
|
@zhfeng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
@zhfeng: This pull request references Jira Issue OCPBUGS-74960, which is valid. 3 validation(s) were run on this bug
Requesting review from QA contact: DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
@openshift-ci-robot: GitHub didn't allow me to request PR reviews from the following users: zhfeng. Note that only openshift members and repo collaborators can review this PR, and authors cannot review their own PRs. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.23-periodics-e2e-aws-private-sg-cleanup |
|
@zhfeng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.23-periodics-e2e-aws-private-sg-cleanup |
|
@zhfeng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.23-periodics-e2e-aws-private-sg-cleanup |
|
@zhfeng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
@zhfeng: This pull request references Jira Issue OCPBUGS-74960, which is valid. 3 validation(s) were run on this bug
Requesting review from QA contact: DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
@openshift-ci-robot: GitHub didn't allow me to request PR reviews from the following users: zhfeng. Note that only openshift members and repo collaborators can review this PR, and authors cannot review their own PRs. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.23-periodics-e2e-aws-private-sg-cleanup |
|
@zhfeng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
@zhfeng: job(s): periodic-ci-openshift-hypershift-release-4.23-periodics-e2e-aws-private-sg-cleanup either don't exist or were not found to be affected, and cannot be rehearsed |
Add CONTROL_PLANE_OPERATOR_IMAGE env var to hypershift-aws-run-e2e-external step to allow overriding the control plane operator image via --e2e.control-plane-operator-image flag. Add periodic job e2e-aws-private-sg-cleanup that runs TestCreateClusterPrivate with a Konflux-built CPO image from openshift/hypershift#7868 to verify the OCPBUGS-74960 fix: security groups for VPC endpoints are properly cleaned up during cluster deletion. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add pre-test step (hypershift-aws-sg-baseline) to record existing vpce-private-router security groups before the test. Add post-test step (hypershift-aws-verify-sg-cleanup) that compares current SGs against the baseline to detect orphaned SGs from the test. Add workflow (hypershift-aws-e2e-external-sg-verify) that wraps the existing external e2e workflow with these SG verification steps. Update the periodic job to use the new sg-verify workflow. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The cli image does not have aws CLI or unzip. Switch to upi-installer which has aws CLI pre-installed.
Instead of flagging all new SGs since baseline (which includes SGs from other parallel tests), check each new SG's VPC for active endpoints. A truly orphaned SG has no active VPC endpoints in its VPC, meaning the endpoint was deleted but the SG was not.
ddaa89c to
eb41924
Compare
|
/pj-rehearse periodic-ci-openshift-hypershift-release-4.23-periodics-e2e-aws-private-sg-cleanup |
|
@zhfeng: now processing your pj-rehearse request. Please allow up to 10 minutes for jobs to trigger or cancel. |
|
[REHEARSALNOTIFIER]
A total of 278 jobs have been affected by this change. The above listing is non-exhaustive and limited to 25 jobs. A full list of affected jobs can be found here Interacting with pj-rehearseComment: Once you are satisfied with the results of the rehearsals, comment: |
|
@zhfeng: all tests passed! Full PR test history. Your PR dashboard. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here. |
|
@zhfeng: This pull request references Jira Issue OCPBUGS-74960. The bug has been updated to no longer refer to the pull request using the external bug tracker. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
Summary
Adds a periodic CI job
e2e-aws-private-sg-cleanupto verify the OCPBUGS-74960 fix prevents VPC endpoint security group leaks when private HostedClusters are deleted.HyperShift PR: openshift/hypershift#7868
Jira: OCPBUGS-74960
What this job does
hypershift-aws-sg-baseline): Records all*-vpce-private-routersecurity groups before the testhypershift-aws-run-e2e-external): RunsTestCreateClusterPrivatewith the CPO image from PR Use namespace resource rather than ProjectRequest to create NS #7868 (via--e2e.control-plane-operator-imageoverride)hypershift-aws-verify-sg-cleanup): After cluster deletion, checks for orphaned SGs by comparing against baseline and verifying each new SG's VPC for active endpointsRehearsal Results (Iteration 4)
SG Verify Output
private-fs4xjSG: properly cleaned up (not in orphaned list)[PASS] All new SGs belong to VPCs with active endpoints/RESULT: ALL CHECKS PASSEDE2e Test Note
The e2e test step failed with a transient
controlPlaneVersion has no desired imagecondition validation timeout (client rate limiter Wait returned an error: context deadline exceeded). The cluster actually deployed successfully (nodes ready in 10m48s, rollout in 5m6s) and was destroyed cleanly. This is a known CI rate-limiter issue unrelated to the fix.CPO Image Override Confirmed
The CPO image from PR #7868 was correctly applied at runtime:
Rehearsal Job
Prow Job Link
New CI Step Registry Components
hypershift-aws-sg-baseline- records SG baseline before testhypershift-aws-verify-sg-cleanup- verifies no orphaned SGs after cluster deletionhypershift-aws-e2e-external-sg-verifyworkflow - orchestrates the full verification/cc @openshift/hypershift-team