Repository navigation
openshift: deployment improvements (keytab migration, Recreate strategy, dist-git SSH fix) - #487
Conversation
There was a problem hiding this comment.
Code Review
This pull request renames the bot keytab to redhat-ymir-agent across various OpenShift configurations, updates the internal SSH configuration, and switches the deployment strategy for several services from RollingUpdate to Recreate. It also adds a helper function to the deployment script to display running pod images and ImageStream tags. Feedback identifies that the new status check in the deployment script may show stale information due to the asynchronous nature of OpenShift deployments and suggests improving the pod output to display human-readable image tags and support multiple containers.
| oc get pods \ | ||
| -o custom-columns='POD:.metadata.name,IMAGE:.status.containerStatuses[0].imageID' \ | ||
| --sort-by='.metadata.name' |
There was a problem hiding this comment.
Using .status.containerStatuses[0].imageID shows the resolved image SHA, which is difficult to verify at a glance (e.g., docker-pullable://...). Additionally, it only shows the first container, which might be incomplete if sidecars are added later. Using .spec.containers[*].image provides the human-readable tags and supports multiple containers. Including the pod status is also helpful to see if pods are still restarting.
| oc get pods \ | |
| -o custom-columns='POD:.metadata.name,IMAGE:.status.containerStatuses[0].imageID' \ | |
| --sort-by='.metadata.name' | |
| oc get pods \ | |
| -o custom-columns='POD:.metadata.name,IMAGE:.spec.containers[*].image,STATUS:.status.phase' \ | |
| --sort-by='.metadata.name' |
| # apply configmap-jira-issue-fetcher-env.yml | ||
| # apply cronjob-jira-issue-fetcher.yml | ||
|
|
||
| show_running_images |
There was a problem hiding this comment.
Since oc apply is asynchronous, calling show_running_images immediately after will likely show the state of the pods before the new deployment has completed. With the Recreate strategy, pods are deleted before new ones are created, so this output might show pods in Terminating state or no pods at all for a brief moment. Consider adding a brief sleep or using oc rollout status for critical deployments to ensure the summary reflects the new state.
There was a problem hiding this comment.
I agree with Gemini, the idea was great but the output is not useful, I'm gonna drop that commit and send a better version in the next PR
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Tomas Tomecek <ttomecek@redhat.com>
Signed-off-by: Tomas Tomecek <ttomecek@redhat.com>
All our deployments run with replicas=1, making RollingUpdate counterproductive: it holds the old pod alive until the new one is ready, which exhausts the namespace memory quota and causes the new pod to fail with FailedCreate. This creates a deadlock that requires manual intervention on every deploy. Recreate simply terminates the old pod first and then starts the new one, which is the right behavior for single-replica workloads. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add explicit requests.memory to all agent deployments (previously unset, defaulting to limits). Bump backport agents from 4Gi to 5Gi limits using the newly available headroom. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Summary
jotnar-bottoredhat-ymir-agentin mcp-gateway deployment and kerberos ConfigMapRollingUpdatetoRecreatestrategy — withreplicas=1and a tight namespace quota, RollingUpdate causes a deadlock where old failing pods hold quota and prevent new pods from being createdUser redhat-ymir-agentto the dist-git SSH config forpkgs.devel.redhat.com(was only set for the bastion)deploy.shTest plan
./openshift/deploy.shruns cleanlyredhat-ymir-agent🤖 Generated with Claude Code