Skip to content

OCPEDGE-2280 OCPEDGE-2640: mutable topology - #2008

Merged
openshift-merge-bot[bot] merged 3 commits into
openshift:masterfrom
jeff-roche:mutable-topology
Aug 13, 2026
Merged

OCPEDGE-2280 OCPEDGE-2640: mutable topology#2008
openshift-merge-bot[bot] merged 3 commits into
openshift:masterfrom
jeff-roche:mutable-topology

Conversation

@jeff-roche

@jeff-roche jeff-roche commented May 11, 2026

Copy link
Copy Markdown
Contributor

Summary

Introduces the Mutable Topology enhancement proposal, which enables OpenShift clusters to transition between topology modes as a Day 2 operation. This replaces the previous Adaptable Topology proposal.

Key Design Decisions

  • Controller in cluster-config-operator (CCO) — A new topology transition controller in CCO watches spec.desiredTopology on the Infrastructure CR, validates preconditions, coordinates the transition across operators, and updates topology status fields when complete. CCO was chosen over CVO, CEO, and MCO (and over a standalone operator) because it owns the config.openshift.io API group and the Infrastructure CR lifecycle. See Alternatives in the proposal for the full placement analysis.
  • No new topology enum values — Transitions move between existing TopologyMode values (SingleReplica, HighlyAvailable, etc.). Operators continue reacting to fixed topology values they already understand. Transition complexity is concentrated in a single controller rather than distributed across 30+ operators.
  • Spec/status contract — Follows the standard Kubernetes pattern: spec.desiredTopology expresses administrator intent; status.controlPlaneTopology reflects observed state. Mirrors the oc adm upgrade pattern (patch spec, controller does the work).
  • Feature-gatedMutableTopology gate progresses through DevPreview → TechPreview → GA. Controller is not registered when the gate is disabled (zero runtime overhead).

Scope

  • Initial transition: SNO → HA compact (3-node) on platform: none
  • CLI: oc adm transition topology HighlyAvailable
  • etcd scaling: CEO handles sequential 1→2→3 member scaling via existing learner-to-voter promotion
  • Failure handling: CEO attempts etcd rollback
  • Upgrade safety: CCO sets Upgradeable=False while a transition is in progress

What Changed (Revision History)

The proposal was revised to base the controller in CCO rather than proposing a dedicated standalone operator (OTTO). Key changes from the prior revision:

  • Controller placement moved from a standalone operator to CCO, with full alternatives analysis (CVO, CEO, MCO, standalone operator, CLI-only)
  • Expanded graduation criteria with per-operator topology dependency matrix requirement
  • Added monitoring/telemetry requirements (Prometheus metrics, alerts) for GA graduation
  • Added Support Procedures section with team ownership, detection, and recovery procedures
  • Clarified etcd scaling risks: the 2-voter intermediate state is unique to Day 2 transitions (does not occur during bootstrapping)
  • Added Upgradeable=False enforcement during transitions to prevent concurrent upgrades

Out of Scope

  • Bidirectional transitions (HA → SNO)
  • HyperShift / hosted control planes
  • MicroShift
  • Automatic node provisioning
  • Cloud platforms (AWS, Azure, GCP) — design does not preclude future support
  • platform: baremetal — pending keepalived resolution

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added support for transitioning single-replica clusters to three-node compact highly available topology on platforms without infrastructure integration.
    • Added an administrator-facing topology setting and asynchronous oc adm transition topology command.
    • Added readiness, quorum, progress, completion, and failure status reporting.
    • Cluster upgrades are blocked during active transitions.
    • Added post-transition health validation and documented recovery behavior for unsuccessful transitions.

@openshift-ci
openshift-ci Bot requested review from bn222 and cooktheryan May 11, 2026 19:46
@jeff-roche jeff-roche changed the title enhancements/topologies: mutable topology enhancement proposal OCPEDGE-2280: mutable topology enhancement proposal May 11, 2026
@openshift-ci-robot openshift-ci-robot added the jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. label May 11, 2026
@openshift-ci-robot

openshift-ci-robot commented May 11, 2026

Copy link
Copy Markdown

@jeff-roche: This pull request references OCPEDGE-2280 which is a valid jira issue.

Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the epic to target either version "5.0." or "openshift-5.0.", but it targets "openshift-4.22" instead.

Details

In response to this:

Summary

  • Introduces the Mutable Topology enhancement, replacing the previous Adaptable Topology proposal
  • Proposes a new optional payload operator (OTTO) to orchestrate topology transitions between existing fixed topology modes, rather than adding a new topology enum
  • Initial scope: SNO to HA compact (3-node) on platform: none

Test plan

  • markdownlint passes (markdownlint-cli2)
  • Reviewer feedback from control plane, API, and architecture teams
  • Template structure validated against guidelines/enhancement_template.md

🤖 Generated with Claude Code

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@jeff-roche jeff-roche changed the title OCPEDGE-2280: mutable topology enhancement proposal OCPEDGE-2280: mutable topology May 11, 2026
@jeff-roche

Copy link
Copy Markdown
Contributor Author

@brandisher brandisher left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm missing a "why" statement covering why a day 2, out-of-payload operator is the right choice for this. The CVO section towards the bottom hints at the why a bit but more explicit detail is needed.

With that in mind, I haven't reviewed the EP fully because I don't understand why this is the approach we're taking. The assessment of CVO seems very light and not enough to exclude that as a potential option to meet the goals.

Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
@jeff-roche

jeff-roche commented May 12, 2026

Copy link
Copy Markdown
Contributor Author

I'm missing a "why" statement covering why a day 2, out-of-payload operator is the right choice for this. The CVO section towards the bottom hints at the why a bit but more explicit detail is needed.

With that in mind, I haven't reviewed the EP fully because I don't understand why this is the approach we're taking. The assessment of CVO seems very light and not enough to exclude that as a potential option to meet the goals.

@brandisher I've added a new paragraph under the ## Proposal header that explains the why. If you're looking for something specifically beyond what I added, I'd be happy to add some more detail

Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated

@JoelSpeed JoelSpeed left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Generated with Claude Code

There are significant portions of this proposal that assume behaviour of OpenShift that either doesn't exist, or doesn't work in the way proposed. I'm assuming here that this is hallucination of Claude?

The EP as it stands today doesn't actually make sense for implementation. It also doesn't align with what I thought we had agreed on the architecture call.

Has anyone tried to manually take a cluster and scale up and manually transition from a single replica to multiple replicas? IMO this is the most important next step for this project

What I thought we had agreed:

  • To scale from SNO to HA, the user must create two new control plane nodes and join them to the cluster
    • On HighlyAvailable topology - KAS, KCM, etcd, etc all get scheduled automatically as static pods on these nodes - I don't see anything that prevents this based on if it's a SNO cluster today, this needs to be checked (it probably should)
    • MCO still serves ignition for control plane nodes on SNO, so user needs to create the control plane nodes somehow to ignite from here
  • New fields are added to the infrastructure spec to allow the user to say "I intend for this cluster to be HA going forward"
  • A controller is added to cluster config operator
    • This checks that the precondition of having additional control plane nodes in the cluster is met
    • Once the precondition is met, it updates the status to reflect spec
  • Operators now react to the change in status and transition from single to HA
    • etcd operator promotes learners to full members, quorum goes from 1->3 (I don't know if this guard is in place today, we should add if not)
    • KAS/KCM - no change, it already scheduled new KAS/kCM pods
    • Others - Those that previously deploy a single replica of their operand now move to 2 replicas, other changes might be needed on a per operator basis, I was expecting those details in the EP but don't see them yet

Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated

@patrickdillon patrickdillon left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know the scope is limited to baremetal/platform:none, but I know there is interest for mutable topologies in cloud platforms as well so as much as appropriate I would to ensure the design leaves a path forward for those cloud platforms.

Also, like the other enhancement I don't see any mention of mastersSchedulable which affects the calculation for infrastructureTopology. How is the mastersSchedulable field handled/taken into account for this solution?

Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated

@zaneb zaneb left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This one looks directionally correct 👍

Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
@jeff-roche

jeff-roche commented May 15, 2026

Copy link
Copy Markdown
Contributor Author

Big update coming next week to realign this with CCO instead of a dedicated operator, add some more technical detail around the flow, and address masters schedulable. Thank you everyone for the quick and thorough reviews, I believe we are rapidly converging on a solid solution!

Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md
@dhensel-rh

Copy link
Copy Markdown
Contributor

Are there limitations for a SNO to TNF transition ? TNF requires BMC/Redfish so if the SNO bare metal hardware does not have it, does it block the transition? I could see this being a problem trying to match hardware in general (BMC firmware versions, vendor types, etc. ).

@jeff-roche jeff-roche changed the title OCPEDGE-2280: mutable topology OCPEDGE-2280: OCPEDGE-2640: mutable topology Aug 6, 2026
@jeff-roche

Copy link
Copy Markdown
Contributor Author

/jira refresh

@openshift-ci-robot

openshift-ci-robot commented Aug 6, 2026

Copy link
Copy Markdown

@jeff-roche: This pull request references OCPEDGE-2280 which is a valid jira issue.

Details

In response to this:

/jira refresh

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@jeff-roche jeff-roche changed the title OCPEDGE-2280: OCPEDGE-2640: mutable topology OCPEDGE-2280 OCPEDGE-2640: mutable topology Aug 6, 2026
@openshift-ci-robot

openshift-ci-robot commented Aug 6, 2026

Copy link
Copy Markdown

@jeff-roche: This pull request references OCPEDGE-2280 which is a valid jira issue.

This pull request references OCPEDGE-2640 which is a valid jira issue.

Details

In response to this:

Summary

Introduces the Mutable Topology enhancement proposal, which enables OpenShift clusters to transition between topology modes as a Day 2 operation. This replaces the previous Adaptable Topology proposal.

Key Design Decisions

  • Controller in cluster-config-operator (CCO) — A new topology transition controller in CCO watches spec.desiredTopology on the Infrastructure CR, validates preconditions, coordinates the transition across operators, and updates topology status fields when complete. CCO was chosen over CVO, CEO, and MCO (and over a standalone operator) because it owns the config.openshift.io API group and the Infrastructure CR lifecycle. See Alternatives in the proposal for the full placement analysis.
  • No new topology enum values — Transitions move between existing TopologyMode values (SingleReplica, HighlyAvailable, etc.). Operators continue reacting to fixed topology values they already understand. Transition complexity is concentrated in a single controller rather than distributed across 30+ operators.
  • Spec/status contract — Follows the standard Kubernetes pattern: spec.desiredTopology expresses administrator intent; status.controlPlaneTopology reflects observed state. Mirrors the oc adm upgrade pattern (patch spec, controller does the work).
  • Feature-gatedMutableTopology gate progresses through DevPreview → TechPreview → GA. Controller is not registered when the gate is disabled (zero runtime overhead).

Scope

  • Initial transition: SNO → HA compact (3-node) on platform: none
  • CLI: oc adm transition topology HighlyAvailable
  • etcd scaling: CEO handles sequential 1→2→3 member scaling via existing learner-to-voter promotion
  • Failure handling: CEO attempts etcd rollback
  • Upgrade safety: CCO sets Upgradeable=False while a transition is in progress

What Changed (Revision History)

The proposal was revised to base the controller in CCO rather than proposing a dedicated standalone operator (OTTO). Key changes from the prior revision:

  • Controller placement moved from a standalone operator to CCO, with full alternatives analysis (CVO, CEO, MCO, standalone operator, CLI-only)
  • Expanded graduation criteria with per-operator topology dependency matrix requirement
  • Added monitoring/telemetry requirements (Prometheus metrics, alerts) for GA graduation
  • Added Support Procedures section with team ownership, detection, and recovery procedures
  • Clarified etcd scaling risks: the 2-voter intermediate state is unique to Day 2 transitions (does not occur during bootstrapping)
  • Added Upgradeable=False enforcement during transitions to prevent concurrent upgrades

Out of Scope

  • Bidirectional transitions (HA → SNO)
  • HyperShift / hosted control planes
  • MicroShift
  • Automatic node provisioning
  • Cloud platforms (AWS, Azure, GCP) — design does not preclude future support
  • platform: baremetal — pending keepalived resolution

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
  • Added support for transitioning single-replica clusters to a three-node compact highly available topology on platforms without infrastructure integration.
  • Added an administrator-facing setting for requesting supported control plane topology transitions.
  • Introduced the asynchronous oc adm transition topology command.
  • Added transition progress, completion, and failure status reporting.
  • Cluster upgrades are blocked during active transitions, with prerequisite validation and recovery handling.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/topologies/mutable-topology.md`:
- Around line 105-109: Define InfrastructureStatus topology fields consistently
as either the converged observed state or the admitted target state. Prefer
retaining observed-state semantics: do not update controlPlaneTopology or
infrastructureTopology until validation and transition completion succeed, and
add a persisted transition phase for in-progress or failed transitions; revise
the affected wording and transition sequence accordingly.
- Around line 301-323: The supported-transition admission predicate must require
the current topology to be SingleReplica with platform none before applying the
listed preconditions. Also explicitly exclude HyperShift and IBI clusters,
either in this checklist or in the supported-transition matcher, so unsupported
three-node clusters cannot be admitted.
- Around line 325-330: The orchestration steps must make admission and the
topology status update conflict-safe: retain the Infrastructure CR
resourceVersion observed during the re-read, require the status update for
controlPlaneTopology and infrastructureTopology to compare against it, and
restart admission when the compare-and-swap conflicts. Do not commit status for
a request whose spec changed after admission.
- Around line 280-284: Clarify the upgrade policy in the MutableTopology
feature-gate lifecycle section: explicitly state whether clusters enabled during
Dev Preview or Tech Preview must move to the Default feature set before standard
upgrades after transition, or revise the policy so they remain on the applicable
no-upgrade feature set. Ensure this guidance is consistent with the stated
feature-set progression and pre-GA upgrade-testing requirement.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: openshift/coderabbit/.coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: a59ababd-20cc-4758-89de-0da6d7e5cd07

📥 Commits

Reviewing files that changed from the base of the PR and between eb958cf and a9bbfb9.

📒 Files selected for processing (1)
  • enhancements/topologies/mutable-topology.md

Comment thread enhancements/topologies/mutable-topology.md
Comment thread enhancements/topologies/mutable-topology.md
Comment thread enhancements/topologies/mutable-topology.md
Comment thread enhancements/topologies/mutable-topology.md
Introduce the Mutable Topology enhancement, which replaces the
previous Adaptable Topology proposal. Instead of a new topology
enum that all operators must interpret, this approach uses a
dedicated operator (OTTO) to orchestrate transitions between
existing fixed topology modes. Initial scope: SNO to HA compact
on platform: none.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

enhancements/topologies: revise mutable topology proposal to base on CCO

Move the topology transition controller from a standalone operator
(OTTO) into cluster-config-operator. CCO owns the config.openshift.io
API group and infrastructure CR lifecycle, making it the natural home.

Key design decisions:
- desiredTopology initialized by installer to match controlPlaneTopology
  (no kubebuilder default — value is cluster-specific)
- Controller triggers on desiredTopology != status.controlPlaneTopology
- On failure, controller resets desiredTopology to current topology
- Upgrade blocked via Upgradeable=False during transitions
- Condition types: TopologyTransitionProgressing, Completed, Failed
- Per-operator topology audit required for Dev Preview entry

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

Add additional reviewers for mutable topology

Finalized reviewers for everyone who participated in the final review.

feat: adding SNO to HA compact pre/transition/post steps

Signed-off-by: Jeff Roche <jeroche@redhat.com>
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md
Comment thread enhancements/topologies/mutable-topology.md
Signed-off-by: Jeff Roche <jeroche@redhat.com>
@atiratree

Copy link
Copy Markdown
Member

LGTM from the Workloads team perspective (cc @ardaguclu) with regard to changes in oc

@JoelSpeed

Copy link
Copy Markdown
Contributor

/lgtm

@openshift-ci openshift-ci Bot added the lgtm Indicates that a PR is ready to be merged. label Aug 12, 2026
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md Outdated
Comment thread enhancements/topologies/mutable-topology.md
@openshift-ci openshift-ci Bot removed the lgtm Indicates that a PR is ready to be merged. label Aug 12, 2026
@Prashanth684 Prashanth684 added the approved Indicates a PR has been approved by an approver from all required OWNERS files. label Aug 13, 2026
@openshift-ci

openshift-ci Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

Approval requirements bypassed by manually added approval.

This pull-request has been approved by:

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

Signed-off-by: Jeff Roche <jeroche@redhat.com>
@openshift-ci

openshift-ci Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

@jeff-roche: all tests passed!

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@jaypoulz

Copy link
Copy Markdown
Contributor

/lgtm

Reapplying final /lgtm post-addressing Prashanth's feedback. Thanks all!

@openshift-ci openshift-ci Bot added the lgtm Indicates that a PR is ready to be merged. label Aug 13, 2026
@openshift-merge-bot
openshift-merge-bot Bot merged commit a1c6069 into openshift:master Aug 13, 2026
3 checks passed
@jeff-roche
jeff-roche deleted the mutable-topology branch August 13, 2026 13:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by an approver from all required OWNERS files. jira/valid-reference Indicates that this PR references a valid Jira ticket of any type. lgtm Indicates that a PR is ready to be merged.

Projects

None yet

Development

Successfully merging this pull request may close these issues.