OSAC-3593: Migrate storage e2e test suite from osac-test-infra (5/5) - #492
Conversation
The feedback controller removes its finalizer before sending the Signal RPC to fulfillment-service. This means the CR disappears from kubectl before the fulfillment-service archives the DB record. Tests that assert immediately after wait_for_deletion hit a race where the UUID is still in the gRPC list. Replace bare assertions with poll_until via wait_for_grpc_removal (up to 60s). In the normal case the UUID is gone on the first poll.
The feedback controller removes its finalizer before sending the Signal RPC to fulfillment-service. This means the CR disappears from kubectl before the fulfillment-service archives the DB record. Tests that assert immediately after wait_for_deletion hit a race where the UUID is still in the gRPC list. Replace bare assertions with poll_until via wait_for_grpc_removal (up to 60s). In the normal case the UUID is gone on the first poll.
…oll-timeouts increase VM runStrategy transition poll timeout
…tion-race NO_ISSUE: fix gRPC deletion race after CR removal
The E2E tests authenticate exclusively via Kubernetes service account tokens, never via Keycloak JWT, and have no multi-tenant isolation or SecurityGroup lifecycle coverage. New test infrastructure: - OsacCLI.get() and get_unchecked() for osac get <resource> - Keycloak JWT token helper (tests/core/keycloak.py) - JWT fixtures: jwt_cli_user, jwt_cli_admin, jwt_grpc_tenant1/2 - GRPCClient: SecurityGroup CRUD, Get for VNet/Subnet/ComputeInstance - K8sClient: SecurityGroup CR queries - Helpers: SecurityGroup wait functions New tests: - JWT List access for all 12 public API resource types x 2 users (24) - Authorization boundary: regular user denied Users, admin allowed (2) - Invalid token rejection (1) - JWT VirtualNetwork lifecycle: create/get/list/delete via JWT (1) - JWT SecurityGroup lifecycle: create/list/delete via JWT (1) - Multi-tenant isolation: tenant1 resource invisible to tenant2 (1) New test file: SecurityGroup lifecycle (SA token): - Create VNet, create SecurityGroup, wait for CR and Ready - Verify Get returns correct name - Delete SecurityGroup and VNet, verify cleanup Extended existing tests with Get assertions: - test_virtual_network_lifecycle: Get after create - test_subnet_lifecycle: Get after create - test_compute_instance_creation: Get after create Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The E2E tests authenticate exclusively via Kubernetes service account tokens, never via Keycloak JWT, and have no multi-tenant isolation or SecurityGroup lifecycle coverage. New test infrastructure: - OsacCLI.get() and get_unchecked() for osac get <resource> - Keycloak JWT token helper (tests/core/keycloak.py) - JWT fixtures: jwt_cli_user, jwt_cli_admin, jwt_grpc_tenant1/2 - GRPCClient: SecurityGroup CRUD, Get for VNet/Subnet/ComputeInstance - K8sClient: SecurityGroup CR queries - Helpers: SecurityGroup wait functions New tests: - JWT List access for all 12 public API resource types x 2 users (24) - Authorization boundary: regular user denied Users, admin allowed (2) - Invalid token rejection (1) - JWT VirtualNetwork lifecycle: create/get/list/delete via JWT (1) - JWT SecurityGroup lifecycle: create/list/delete via JWT (1) - Multi-tenant isolation: tenant1 resource invisible to tenant2 (1) New test file: SecurityGroup lifecycle (SA token): - Create VNet, create SecurityGroup, wait for CR and Ready - Verify Get returns correct name - Delete SecurityGroup and VNet, verify cleanup Extended existing tests with Get assertions: - test_virtual_network_lifecycle: Get after create - test_subnet_lifecycle: Get after create - test_compute_instance_creation: Get after create Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…ke-tests NO-ISSUE: bring VMaaS E2E test suite up to date
…ke-tests NO-ISSUE: bring VMaaS E2E test suite up to date
Add test-caas Makefile target, kubeconfig/password retrieval tests, template immutability test, and fix cluster grpc removal to use polling instead of one-shot assertion. Update default template to ocp_ci_small for CI environments.
Add test-caas Makefile target, kubeconfig/password retrieval tests, template immutability test, and fix cluster grpc removal to use polling instead of one-shot assertion. Update default template to ocp_ci_small for CI environments.
…to credentials teardown
Co-authored-by: Akshay Nadkarni <25892229+akshaynadkarni@users.noreply.github.com>
Co-authored-by: Akshay Nadkarni <25892229+akshaynadkarni@users.noreply.github.com>
Co-authored-by: Akshay Nadkarni <25892229+akshaynadkarni@users.noreply.github.com>
MGMT-22635: add CaaS test coverage and test-caas target
MGMT-22635: add CaaS test coverage and test-caas target
The `\!=` causes a SyntaxError preventing pytest from collecting the test file. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…utability-syntax fix: remove backslash escaping in assert statement
When an AAP deprovision job fails (intermittent receptor worker stream drop, ~2-5% rate), the operator retries with exponential backoff. The retry succeeds within ~7 minutes but the previous 300s (5 min) timeout on kubectl delete and deletion wait polls expires before the retry completes, failing the test. Bump all deletion waits to 120 retries x 5s = 600s (10 min) and kubectl delete timeout to 600s to accommodate the operator retry.
…imeout NO_ISSUE: bump deletion timeouts from 300s to 600s
wait_for_provision and wait_for_running poll until success but do not check for terminal failure. When provisioning fails immediately (e.g. DataVolumeError), each test wastes 10-15 minutes of the 60-minute CI timeout before raising TimeoutError. Assert state/phase != Failed so the test fails immediately with a clear error instead of burning the timeout budget. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…-provision-wait NO_ISSUE: helpers - fail fast when provision job or phase enters Failed
…sion Previously wait_for_provision checked the latest AAP job state and failed immediately if it saw "Failed". The operator retries failed provision jobs with exponential backoff, so a transient AAP failure followed by a successful retry is normal behavior. Asserting on individual job state made the test fragile and caused false failures in CI. The fix polls the Provisioned status condition instead, which reflects the end result (infrastructure provisioned) regardless of how many AAP job attempts it took. Signed-off-by: akshaynadkarni <25892229+akshaynadkarni@users.noreply.github.com> Assisted-by: Cursor/Claude Signed-off-by: akshaynadkarni <25892229+akshaynadkarni@users.noreply.github.com>
Add a phase check so wait_for_provision fails immediately if the ComputeInstance reaches Failed phase, rather than waiting the full timeout. Transient AAP job failures (which keep the phase at Starting) are still tolerated until the Provisioned condition becomes True. Signed-off-by: akshaynadkarni <25892229+akshaynadkarni@users.noreply.github.com> Assisted-by: Cursor/Claude Signed-off-by: akshaynadkarni <25892229+akshaynadkarni@users.noreply.github.com>
…rovision-condition fix: use Provisioned condition instead of job state in wait_for_provision
…ovision test Replace _wait_for_provision_job (polls internal AAP job ID) and the single-read job state assertion with a poll for phase == "Starting". The CR phase is the authoritative user-facing signal that provisioning is in progress. The AAP job ID is internal bookkeeping. Mirrors the same fix applied to the CaaS equivalent in PR osac-project#50. Co-Authored-By: Claude <noreply@anthropic.com>
…tion bug - test_cluster_explicit_fields: use template parameters (-p/-f) instead of deprecated --pull-secret-file/--ssh-public-key-file CLI flags, update assertions to check templateParameters instead of top-level spec fields - test_cluster_order_delete_during_provision: check CR phase (Progressing) instead of internal job state which races with operator retries, replace bare grpc assert with retried wait_for_cluster_grpc_removal - helpers: work around hypershift bug where capi-provider-agent controller is killed during HostedCluster teardown before it can remove the AgentCluster deprovision finalizer, causing infinite deletion deadlock. Force-remove orphaned finalizers during wait_for_cluster_deletion poll. Co-Authored-By: Claude Code <noreply@anthropic.com>
…tion bug - test_cluster_explicit_fields: use template parameters (-p/-f) instead of deprecated --pull-secret-file/--ssh-public-key-file CLI flags, update assertions to check templateParameters instead of top-level spec fields - test_cluster_order_delete_during_provision: check CR phase (Progressing) instead of internal job state which races with operator retries, replace bare grpc assert with retried wait_for_cluster_grpc_removal - helpers: work around hypershift bug where capi-provider-agent controller is killed during HostedCluster teardown before it can remove the AgentCluster deprovision finalizer, causing infinite deletion deadlock. Force-remove orphaned finalizers during wait_for_cluster_deletion poll. Co-Authored-By: Claude Code <noreply@anthropic.com>
…-fragile-assertions NO_ISSUE: use CR phase instead of job state in VMaaS delete-during-provision test
|
/approve |
|
/hold cancel |
|
@amej: This pull request references OSAC-3593 which is a valid jira issue. Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the task to target the "5.1.0" version, but no target version was set. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
Pull request was closed
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: akshaynadkarni, amej, ronniel1 The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
95c2af4
Address automated review findings on storage e2e test files from PR osac-project#492: - Extract duplicated K8s manifest templates into shared conftest constants - Wrap teardown steps individually to prevent cascading cleanup failures - Let AssertionError propagate from verification steps so they fail the test - Move namespace cleanup out of _verify_teardown into the finally block in test_tenant_storage_lifecycle to prevent namespace leaks - Normalize poll_until checked=False in poll lambdas for consistent error handling - Always verify ClusterOrder removal even on fast deletion path - Assert tenant-scoped secrets are cleaned up during teardown - Add docstrings to all functions and fixtures (coverage improvement from ~7.69%) Signed-off-by: aipcc-bot <aipcc-bot@redhat.com> Co-authored-by: aipcc-bot <aipcc-bot@redhat.com>
Summary
History-preserving migration of the storage e2e test suite from osac-test-infra into this mono-repo. This is PR 5 of 5 — the final PR in a stacked series(bmaas → vmaas → caas → catalog → storage).
Commit history from osac-test-infra is preserved via git filter-repo,so git log tests/storage/ shows the original authors and messages.
What this PR adds
Stack overview
With this PR, all 5 test suites from osac-test-infra are migrated. Remaining OSAC-3593 work: update the e2e workflow to run pytest from this repo, validate with a live e2e run, then remove tests/ from osac-test-infra.
Context
Test plan
Summary by CodeRabbit