Skip to content

Enable TLS encryption in transit for all Ymir services - #864

Open
majamassarini wants to merge 8 commits into
packit:mainfrom
majamassarini:tls-compliance-remediation
Open

majamassarini wants to merge 8 commits into
packit:mainfrom
majamassarini:tls-compliance-remediation

Conversation

@majamassarini

Copy link
Copy Markdown
Member

Summary

  • HSTS headers on all OpenShift Routes with HTTP-to-HTTPS redirect
  • TLS on PostgreSQL (phoenix-db) via OpenShift service-ca, with sslmode=verify-full on the Phoenix client
  • TLS on Valkey — TLS-only port, client auth disabled, all consumers use rediss:// with CA verification
  • TLS on trace-server — reencrypt Routes for classified span data
  • Service-CA bundle mounted into all deployments and CronJobs that need CA trust
  • Annual rotation procedure documented with step-by-step runbook for proactive certificate rotation

Post-merge action: create recurring Jira ticket

Once this PR is merged, create a recurring Jira issue with:

  • Summary: Annual TLS certificate rotation — Ymir (jotnar-ymir)
  • Type: Task
  • Priority: High
  • Due date: 12 months from merge date
  • Description:
h2. Purpose

Proactively rotate OpenShift service-CA serving certificates and restart
all workloads before the 13-month grace period expires.

See: docs/tls-compliance.md § "Annual Rotation Procedure"

h2. Steps

1. *Rotate serving certificates*
{code}
oc delete secret phoenix-db-tls valkey-tls otel-collector-tls
{code}

2. *Restart TLS endpoints (servers first)*
{code}
oc rollout restart deployment/phoenix-db deployment/valkey deployment/otel-collector
oc rollout status deployment/phoenix-db
oc rollout status deployment/valkey
oc rollout status deployment/otel-collector
{code}

3. *Restart all client deployments*
{code}
oc rollout restart deployment/api
oc rollout restart deployment/phoenix
oc rollout restart deployment/redis-commander
oc rollout restart deployment/backport-agent-c9s deployment/backport-agent-c10s
oc rollout restart deployment/rebase-agent-c9s deployment/rebase-agent-c10s
oc rollout restart deployment/rebuild-agent-c9s deployment/rebuild-agent-c10s
oc rollout restart deployment/mr-consolidation-agent-c9s deployment/mr-consolidation-agent-c10s
oc rollout restart deployment/triage-agent deployment/reproducer-agent
{code}

4. *Verify*
{code}
for secret in phoenix-db-tls valkey-tls otel-collector-tls; do
  oc get secret "$secret" \
    -o jsonpath="${secret}: expiry={.metadata.annotations.service\.beta\.openshift\.io/expiry} version={.metadata.resourceVersion}{'\n'}"
done
{code}
Confirm all deployments are Ready and no TLS errors in logs.

5. *Clone this ticket* with due date = today + 12 months.

h2. Acceptance criteria

* [ ] All three secrets regenerated with new expiry dates
* [ ] All deployments restarted and Ready
* [ ] No TLS errors in application logs
* [ ] Next year's rotation ticket created

Test plan

  • Deploy to staging and verify all Routes return Strict-Transport-Security header
  • Confirm HTTP requests redirect to HTTPS on all Routes
  • Verify PostgreSQL SSL: oc exec deployment/phoenix-db -- psql -c "SHOW ssl" returns on
  • Verify Valkey TLS-only: oc exec deployment/valkey -- valkey-cli --tls --cert /tls/tls.crt --key /tls/tls.key --cacert /etc/pki/service-ca/service-ca.crt ping
  • Verify trace-server Routes use reencrypt termination
  • Confirm all agent pods start successfully with rediss:// URLs
  • Run the annual rotation procedure on staging to validate the runbook

🤖 Generated with Claude Code

@qodo-for-packit

Copy link
Copy Markdown

PR Summary by Qodo

Enable TLS across Ymir database, cache, and trace services

✨ Enhancement ⚙️ Configuration changes 📝 Documentation 🕐 40+ Minutes

Grey Divider

AI Description

• Encrypt PostgreSQL and Valkey connections using OpenShift-issued certificates and service-CA
 trust.
• Add HSTS to all Routes and reencrypt traffic from the router to trace-server.
• Document TLS verification and annual certificate rotation, including required workload restarts.
Diagram

graph TD
  CA["OpenShift service CA"] --> Bundle["CA trust bundle"] --> Phoenix["Phoenix client"] --> DB[("PostgreSQL")]
  Bundle --> Agents["Agents and tools"] --> Valkey[("Valkey")]
  CA --> DB
  CA --> Valkey
  CA --> Trace["Trace server"]
  Routes["Trace Routes"] --> Trace
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Automate certificate-triggered workload restarts
  • ➕ Removes dependence on an annual ticket to refresh certificates loaded at startup.
  • ➖ Adds rollout automation and operational complexity across stateful services and clients.

Recommendation: Using OpenShift service certificates and its CA bundle fits the existing deployment model. Keep the documented annual procedure for this change, but consider restart automation if manual rotation becomes an operational risk.

Files changed (40) +493 / -18

Enhancement (2) +17 / -2
server.pySupport HTTPS in trace-server +11/-1

Support HTTPS in trace-server

• Wraps the HTTP server socket with TLS when certificate and key paths are configured; retains the existing non-TLS startup path otherwise.

trace_server/server.py

base_utils.pySupply CA trust to Redis TLS clients +6/-1

Supply CA trust to Redis TLS clients

• Passes a configured or default CA file to redis-py for rediss URLs when that file exists.

ymir/common/base_utils.py

Documentation (1) +216 / -0
tls-compliance.mdDocument encryption posture and rotation +216/-0

Document encryption posture and rotation

• Describes Route, database, cache, and trace-server TLS controls, plus verification and annual certificate-rotation steps.

docs/tls-compliance.md

Other (37) +260 / -16
.secrets.baselineRefresh secret-scan locations +10/-10

Refresh secret-scan locations

• Adjusts recorded line numbers after manifest edits and refreshes the baseline timestamp.

.secrets.baseline

configmap-endpoints-env.ymlSwitch shared Valkey URL to TLS +1/-1

Switch shared Valkey URL to TLS

• Changes the shared Redis endpoint to the rediss scheme.

openshift/configmap-endpoints-env.yml

configmap-otel-collector-config.ymlExport traces to the HTTPS sidecar +3/-1

Export traces to the HTTPS sidecar

• Changes the local trace-server exporter endpoint to HTTPS and skips certificate verification for that localhost connection.

openshift/configmap-otel-collector-config.yml

configmap-service-ca.ymlRequest an injected service-CA bundle +7/-0

Request an injected service-CA bundle

• Adds a ConfigMap annotated for OpenShift service-CA injection.

openshift/configmap-service-ca.yml

cronjob-jira-issue-fetcher-todo.ymlMount CA trust in TODO issue fetcher +8/-0

Mount CA trust in TODO issue fetcher

• Mounts the service-CA bundle read-only in the CronJob pod.

openshift/cronjob-jira-issue-fetcher-todo.yml

cronjob-jira-issue-fetcher.ymlMount CA trust in issue fetcher +8/-0

Mount CA trust in issue fetcher

• Mounts the service-CA bundle read-only in the CronJob pod.

openshift/cronjob-jira-issue-fetcher.yml

cronjob-supervisor-collector.ymlMount CA trust in supervisor collector +6/-0

Mount CA trust in supervisor collector

• Adds a read-only service-CA bundle mount to the collector CronJob.

openshift/cronjob-supervisor-collector.yml

cronjob-sweep-dependency.ymlMount CA trust in dependency sweep +6/-0

Mount CA trust in dependency sweep

• Adds a read-only service-CA bundle mount to the sweep CronJob.

openshift/cronjob-sweep-dependency.yml

cronjob-sweep-no-patch.ymlMount CA trust in no-patch sweep +8/-0

Mount CA trust in no-patch sweep

• Adds a read-only service-CA bundle mount to the sweep CronJob.

openshift/cronjob-sweep-no-patch.yml

cronjob-sweep-pr-pending.ymlMount CA trust in pending-PR sweep +8/-0

Mount CA trust in pending-PR sweep

• Adds a read-only service-CA bundle mount to the sweep CronJob.

openshift/cronjob-sweep-pr-pending.yml

cronjob-sweep-y-stream.ymlMount CA trust in Y-stream sweep +6/-0

Mount CA trust in Y-stream sweep

• Adds a read-only service-CA bundle mount to the sweep CronJob.

openshift/cronjob-sweep-y-stream.yml

deploy-oc.shApply CA bundle before dependent workloads +3/-0

Apply CA bundle before dependent workloads

• Creates the service-CA ConfigMap early in the OpenShift deployment sequence.

openshift/deploy-oc.sh

deployment-api.ymlMount CA trust in API +8/-0

Mount CA trust in API

• Adds a read-only service-CA bundle mount to the API deployment.

openshift/deployment-api.yml

deployment-backport-agent-c10s.ymlMount CA trust in c10s backport agent +6/-0

Mount CA trust in c10s backport agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-backport-agent-c10s.yml

deployment-backport-agent-c9s.ymlMount CA trust in c9s backport agent +6/-0

Mount CA trust in c9s backport agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-backport-agent-c9s.yml

deployment-mr-consolidation-agent-c10s.ymlMount CA trust in c10s consolidation agent +6/-0

Mount CA trust in c10s consolidation agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-mr-consolidation-agent-c10s.yml

deployment-mr-consolidation-agent-c9s.ymlMount CA trust in c9s consolidation agent +6/-0

Mount CA trust in c9s consolidation agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-mr-consolidation-agent-c9s.yml

deployment-otel-collector.ymlServe trace-server HTTPS +13/-0

Serve trace-server HTTPS

• Mounts a serving certificate in the trace-server sidecar, configures its TLS paths, and changes its health probes to HTTPS.

openshift/deployment-otel-collector.yml

deployment-phoenix-db.ymlConfigure PostgreSQL serving TLS +38/-0

Configure PostgreSQL serving TLS

• Adds an init container to prepare the serving certificate, key permissions, and PostgreSQL SSL configuration; mounts the prepared files into the database.

openshift/deployment-phoenix-db.yml

deployment-phoenix.ymlVerify PostgreSQL certificate from Phoenix +7/-1

Verify PostgreSQL certificate from Phoenix

• Requires verify-full TLS for the database connection and mounts the service-CA bundle as its trust root.

openshift/deployment-phoenix.yml

deployment-rebase-agent-c10s.ymlMount CA trust in c10s rebase agent +6/-0

Mount CA trust in c10s rebase agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-rebase-agent-c10s.yml

deployment-rebase-agent-c9s.ymlMount CA trust in c9s rebase agent +6/-0

Mount CA trust in c9s rebase agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-rebase-agent-c9s.yml

deployment-rebuild-agent-c10s.ymlMount CA trust in c10s rebuild agent +6/-0

Mount CA trust in c10s rebuild agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-rebuild-agent-c10s.yml

deployment-rebuild-agent-c9s.ymlMount CA trust in c9s rebuild agent +6/-0

Mount CA trust in c9s rebuild agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-rebuild-agent-c9s.yml

deployment-redis-commander.ymlConfigure Redis Commander for TLS +10/-0

Configure Redis Commander for TLS

• Sets TLS-related environment variables and mounts the service-CA bundle for Valkey access.

openshift/deployment-redis-commander.yml

deployment-reproducer-agent.ymlMount CA trust in reproducer agent +6/-0

Mount CA trust in reproducer agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-reproducer-agent.yml

deployment-supervisor-processor.ymlMount CA trust in supervisor processor +6/-0

Mount CA trust in supervisor processor

• Adds a read-only service-CA bundle mount to the deployment.

openshift/deployment-supervisor-processor.yml

deployment-triage-agent.ymlMount CA trust in triage agent +6/-0

Mount CA trust in triage agent

• Adds a read-only service-CA bundle mount to the agent.

openshift/deployment-triage-agent.yml

deployment-valkey.ymlServe Valkey over TLS only +26/-1

Serve Valkey over TLS only

• Disables the plaintext port, configures TLS without client-certificate authentication, and mounts the serving certificate and CA bundle.

openshift/deployment-valkey.yml

route-api.ymlAdd HSTS to API Route +2/-0

Add HSTS to API Route

• Sets a one-year HSTS header annotation on the API Route.

openshift/route-api.yml

route-phoenix.ymlAdd HSTS to Phoenix Route +2/-0

Add HSTS to Phoenix Route

• Sets a one-year HSTS header annotation on the Phoenix Route.

openshift/route-phoenix.yml

route-redis-commander.ymlAdd HSTS to Redis Commander Route +2/-0

Add HSTS to Redis Commander Route

• Sets a one-year HSTS header annotation on the Redis Commander Route.

openshift/route-redis-commander.yml

route-trace-server-cname.ymlReencrypt custom trace Route +3/-1

Reencrypt custom trace Route

• Adds HSTS and changes the custom-domain trace Route from edge termination to reencrypt termination.

openshift/route-trace-server-cname.yml

route-trace-server.ymlReencrypt trace-server Route +3/-1

Reencrypt trace-server Route

• Adds HSTS and changes the trace-server Route from edge termination to reencrypt termination.

openshift/route-trace-server.yml

service-otel-collector.ymlRequest trace-server serving certificate +2/-0

Request trace-server serving certificate

• Annotates the collector Service to request an OpenShift-issued serving certificate.

openshift/service-otel-collector.yml

service-phoenix-db.ymlRequest PostgreSQL serving certificate +2/-0

Request PostgreSQL serving certificate

• Annotates the database Service to request an OpenShift-issued serving certificate.

openshift/service-phoenix-db.yml

service-valkey.ymlRequest Valkey serving certificate +2/-0

Request Valkey serving certificate

• Annotates the Valkey Service to request an OpenShift-issued serving certificate.

openshift/service-valkey.yml

@majamassarini
majamassarini force-pushed the tls-compliance-remediation branch from 6c0297b to 28b8453 Compare October 8, 2026 12:51
@qodo-for-packit

qodo-for-packit Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (6) 📘 Rule violations (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Redis clients cannot complete the handshake ✓ Resolved
Description
The Valkey TLS configuration supplies a CA but does not disable TLS client authentication, while the
Python clients and Redis Commander supply no client certificate or key. When Valkey requests a
client certificate, neither class of consumer can connect to its TLS-only port.
Code

openshift/deployment-valkey.yml[R28-31]

- --tls-key-file
- /tls/tls.key
- --tls-ca-cert-file
- /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt
Relevance

●●● Strong

TLS-only Valkey with default client authentication conflicts directly with clients lacking
certificates; this is a concrete handshake failure.

PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The Valkey command has no client-auth override; the shared Python helper has no client-certificate
parameters, and the Commander deployment provides only a CA path.

openshift/deployment-valkey.yml[21-31]
ymir/common/base_utils.py[55-64]
openshift/deployment-redis-commander.yml[39-69]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Valkey requires TLS client authentication but its consumers have no client credentials.
## Fix Focus Areas
- openshift/deployment-valkey.yml[21-31]
- openshift/deployment-redis-commander.yml[45-48]
- ymir/common/base_utils.py[55-64]
## Recommended Fix
Either explicitly disable Valkey TLS client authentication or provision and configure trusted client certificates and keys for every consumer.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Phoenix cannot reach its database ✓ Resolved
Description
PHOENIX_SQL_DATABASE_URL connects to the short host phoenix-db with sslmode=verify-full, but
the phoenix-db-tls serving certificate covers only phoenix-db.<namespace>.svc and
phoenix-db.<namespace>.svc.cluster.local. When Phoenix validates that certificate at startup,
hostname verification rejects the database connection, affecting the trace UI, trace storage, and
OTel → Phoenix export.
Code

openshift/deployment-phoenix.yml[53]

+          value: postgresql://$(PHOENIX_POSTGRES_USER):$(PH**********************D)@phoenix-db/$(PHOENIX_POSTGRES_DB)?sslmode=verify-full&sslrootcert=/etc/pki/service-ca/service-ca.crt
Relevance

●●● Strong

Using verify-full with an incompatible short service hostname is a concrete startup-breaking TLS
configuration defect.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The Phoenix URL uses phoenix-db as its host and enables verify-full. The Service requests an
OpenShift serving certificate through the serving-cert-secret-name: phoenix-db-tls annotation, and
the DB deployment copies that certificate into PostgreSQL; the certificate identifies the service by
its namespace-qualified DNS names rather than the short alias. Because the deployment namespace
makes the .svc name phoenix-db.jotnar-ymir--jotnar-ymir.svc, the URL's short host fails the
verify-full hostname check.

openshift/deployment-phoenix.yml[52-53]
openshift/service-phoenix-db.yml[1-8]
openshift/service-phoenix-db.yml[3-6]
openshift/deployment-phoenix-db.yml[44-51]
openshift/deployment-phoenix.yml[50-53]
openshift/deploy-oc.sh[5-8]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Phoenix connects to the short host `phoenix-db` with `sslmode=verify-full`, but the service-ca serving certificate identifies the database by its namespace-qualified service DNS name, so hostname verification fails.

## Fix Focus Areas
- openshift/deployment-phoenix.yml[50-53]
- openshift/service-phoenix-db.yml[3-6]

## Recommended Fix
Change the host in `PHOENIX_SQL_DATABASE_URL` to `phoenix-db.jotnar-ymir--jotnar-ymir.svc`, retaining `sslmode=verify-full&sslrootcert=/etc/pki/service-ca/service-ca.crt`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. All agents and cronjobs lose access to the queue ✓ Resolved
Description
REDIS_URL now points to rediss://valkey:6379/0. Every Python consumer connects through
redis_client, which uses redis-py >= 6.4, and redis-py 6.x checks the hostname by default
(ssl_check_hostname=True). OpenShift service-ca serving certificates only cover
valkey.<namespace>.svc and valkey.<namespace>.svc.cluster.local, not the short name valkey. So
once Valkey switches to TLS-only, every TLS handshake from agents, API, supervisor, fetcher and
sweep jobs fails on a hostname mismatch, and redis-commander (REDIS_HOST=valkey with
REDIS_TLS=true) fails the same way.
Code

openshift/configmap-endpoints-env.yml[7]

+  REDIS_URL: rediss://valkey:6379/0
Relevance

●●● Strong

TLS hostname mismatch would break every Redis consumer, directly contradicting the PR’s stated TLS
objective.

PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Both requirements files pin redis>=6.4.0, a version where hostname checking is on by default. All
Redis clients go through redis_client → redis.Redis.from_url(redis_url, ...), which only adds
ssl_ca_certs and never disables or overrides the hostname check. The URL host is the short Service
name valkey, but the certificate comes from the
service.beta.openshift.io/serving-cert-secret-name: valkey-tls annotation on the valkey Service,
and those certificates carry only the .svc and .svc.cluster.local names. redis-commander also
connects by the short name.

ymir/common/base_utils.py[63-68]
requirements.txt[15-15]
openshift/service-valkey.yml[1-8]
openshift/deployment-redis-commander.yml[41-48]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Redis clients connect to `rediss://valkey:6379/0`, but the service-ca certificate only lists `valkey.<ns>.svc` and `valkey.<ns>.svc.cluster.local`. redis-py 6.x checks the hostname by default, so every TLS handshake fails.

## Fix Focus Areas
- openshift/configmap-endpoints-env.yml[7-7]
- openshift/deployment-redis-commander.yml[41-48]

## Recommended Fix
Change `REDIS_URL` to `rediss://valkey.jotnar-ymir--jotnar-ymir.svc:6379/0`, or use the `.svc.cluster.local` form. Set redis-commander's `REDIS_HOST` to the same full name. Then confirm with `valkey-cli --tls --cacert ... -h valkey.<ns>.svc ping` from a client pod.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View high (4)
4. Certificate checks never get deployed 🐞 Bug ≡ Correctness 👈 Shift left
Description
deploy-oc.sh applies the existing sweep CronJobs but never applies the new tls-cert-check
manifest. Deployments made through that script therefore create no weekly certificate-check Job,
regardless of whether the checker image is available.
Code

openshift/cronjob-tls-cert-check.yml[R1-4]

apiVersion: batch/v1
kind: CronJob
metadata:
name: tls-cert-check
Relevance

●●● Strong

Deployment omission is a deterministic correctness bug preventing the newly introduced CronJob from
existing.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The repository's deployment script explicitly applies the four existing sweep CronJobs after
importing their image; its list has no entry for the newly defined certificate checker.

openshift/deploy-oc.sh[157-165]
openshift/cronjob-tls-cert-check.yml[1-9]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The new certificate-check CronJob is defined but omitted from the deployment script's explicit manifest list.
## Fix Focus Areas
- openshift/cronjob-tls-cert-check.yml[1-4]
- openshift/deploy-oc.sh[157-165]
## Recommended Fix
Add an `apply cronjob-tls-cert-check.yml` command to the deployment script after the sweep ImageStream is imported, so normal deployments create and update the CronJob.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Dismiss ↗ | View ↗


5. Valkey cannot load its configured TLS CA ✓ Resolved
Description
deployment-valkey.yml passes a service-CA path to --tls-ca-cert-file without mounting a file at
that path. When the TLS-only Valkey process starts, its pod supplies only the data volume and the
serving-certificate secret, so the configured CA cannot be loaded.
Code

openshift/deployment-valkey.yml[R30-31]

- --tls-ca-cert-file
- /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt
Relevance

●●● Strong

Valkey is configured for TLS but its CA path is not mounted, causing deterministic process startup
failure.

PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The added argument names a file outside both configured mounts, and the deployment script applies no
service-CA ConfigMap.

openshift/deployment-valkey.yml[21-31]
openshift/deployment-valkey.yml[46-66]
openshift/deploy-oc.sh[99-107]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Valkey references a service-CA file that its pod does not provide.
## Fix Focus Areas
- openshift/deployment-valkey.yml[21-66]
- openshift/deploy-oc.sh[99-107]
## Recommended Fix
Create an OpenShift service-CA-injected ConfigMap, mount its bundle into Valkey, and point `--tls-ca-cert-file` at the mounted file.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


6. Redis Commander cannot load its CA ✓ Resolved
Description
REDIS_TLS_CA_CERT_FILE points to a service-CA file absent from the Redis Commander container. Its
deployment mounts only the configuration volume, so the new TLS connection cannot use the specified
CA certificate.
Code

openshift/deployment-redis-commander.yml[R47-48]

- name: REDIS_TLS_CA_CERT_FILE
value: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt
Relevance

●●● Strong

The configured CA path is absent from the pod, making the newly enabled Redis TLS connection
unusable.

PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The newly configured CA path is not supplied by the container's sole config mount.

openshift/deployment-redis-commander.yml[41-48]
openshift/deployment-redis-commander.yml[58-69]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Redis Commander is given a TLS CA path without the corresponding file.
## Fix Focus Areas
- openshift/deployment-redis-commander.yml[41-69]
## Recommended Fix
Mount an injected service-CA bundle in the Redis Commander container and update `REDIS_TLS_CA_CERT_FILE` to its mount path.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


7. Database access fails after cert rotation 🐞 Bug ☼ Reliability 👈 Shift left
Description
setup-tls copies the PostgreSQL serving certificate and key into emptyDir only when the pod
starts. If the service-CA operator rotates the Secret while that pod remains running, PostgreSQL
retains the old files and Phoenix's verify-full connections eventually reject the expired
certificate.
Code

openshift/deployment-phoenix-db.yml[R45-46]

cp /tls-secret/tls.crt /pg-tls/server.crt
cp /tls-secret/tls.key /pg-tls/server.key
Relevance

●●● Strong

Init-only certificate copying creates a concrete rotation failure, matching the team’s acceptance of
deployment reliability fixes.

PR-#703
PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The Service requests an operator-managed serving Secret, but the database reads files copied by a
terminating init container, with no subsequent update mechanism in the deployment.

openshift/service-phoenix-db.yml[3-6]
openshift/deployment-phoenix-db.yml[31-51]
openshift/deployment-phoenix-db.yml[91-115]
openshift/deployment-phoenix.yml[53-53]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
PostgreSQL serves a one-time copy of a rotating serving certificate.
## Fix Focus Areas
- openshift/deployment-phoenix-db.yml[31-51]
- openshift/deployment-phoenix-db.yml[91-115]
## Recommended Fix
Add a certificate-rotation mechanism that updates PostgreSQL's usable files and reloads or restarts PostgreSQL when the serving Secret changes.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Dismiss ↗ | View ↗



Remediation recommended

8. Redis Commander likely can't reach the cache over TLS ✗ Dismissed
Description
The redis-commander image starts with `redis-commander --redis-host $REDIS_HOST --redis-port
$REDIS_PORT and no --redis-tls flag, so the new REDIS_TLS / REDIS_TLS_CA_CERT_FILE` env vars
are probably never used: that mapping is done by the upstream image's entrypoint script, which this
custom Fedora image doesn't run. Valkey is now TLS-only (--port 0), so the Redis Commander UI
would send plaintext to the TLS port and fail to connect.
Code

openshift/deployment-redis-commander.yml[R45-48]

+        - name: REDIS_TLS
+          value: "true"
+        - name: REDIS_TLS_CA_CERT_FILE
+          value: /etc/pki/service-ca/service-ca.crt
Relevance

●●● Strong

This is a concrete connectivity bug; accepted precedents favor correcting command-line and
deployment integration failures.

PR-#584
PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The image is built from Containerfile.redis-commander, a plain npm install -g redis-commander on
Fedora. Its CMD passes only host and port, with no TLS option and no entrypoint that turns env vars
into config. Valkey now listens only on its TLS port, so a plaintext client can't connect.

Containerfile.redis-commander[1-13]
openshift/deployment-redis-commander.yml[40-48]
openshift/deployment-valkey.yml[21-33]
openshift/imagestream-redis-commander.yml[10-10]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
redis-commander is started via a custom CMD that only passes `--redis-host`/`--redis-port`, so the new `REDIS_TLS` / `REDIS_TLS_CA_CERT_FILE` env vars are probably ignored and the UI can't connect to TLS-only Valkey.

## Fix Focus Areas
- Containerfile.redis-commander[13-13]
- openshift/deployment-redis-commander.yml[40-48]

## Recommended Fix
Add `--redis-tls` to the CMD (for example, make it conditional on `$REDIS_TLS`). Configure the CA by writing a `local-production.json` into the copied config dir, with `redis.tls.ca` set to the service CA file contents and `servername` set to `$REDIS_HOST`. Alternatively, switch to an entrypoint that reads these env vars. Then confirm on staging that the UI lists keys.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


9. Fresh deployments can start without CA trust 🐞 Bug ☼ Reliability ⭐ New 👈 Shift left
Description
deploy-oc.sh applies service-ca-bundle but does not wait for the injected service-ca.crt
before deploying Valkey and its clients. On a fresh deployment where CA injection has not finished,
Valkey starts with a path to a missing CA file and TLS clients cannot establish trusted connections
until the bundle arrives and they recover.
Code

openshift/deploy-oc.sh[R68-69]

# Service-CA bundle (must be applied before TLS-consuming deployments)
apply configmap-service-ca.yml
Relevance

●●● Strong

Fresh TLS deployments can race asynchronous CA injection; reviewers accept concrete OpenShift
startup reliability fixes.

PR-#443
PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new ConfigMap starts with empty data, the deployment script proceeds without checking its
contents, and Valkey's startup arguments require the certificate at the mounted path. The Redis
client likewise only supplies that CA when the file exists.

openshift/configmap-service-ca.yml[3-7]
openshift/deploy-oc.sh[68-70]
openshift/deploy-oc.sh[96-101]
openshift/deployment-valkey.yml[26-31]
ymir/common/base_utils.py[63-68]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Applying the initially empty service-CA ConfigMap does not ensure its certificate has been injected before TLS workloads start.

## Fix Focus Areas
- openshift/deploy-oc.sh[68-70]
- openshift/configmap-service-ca.yml[3-7]

## Recommended Fix
After applying the ConfigMap, poll until `service-ca.crt` is populated, with a bounded timeout and a clear deployment error, before applying TLS-consuming workloads.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Dismiss ↗ | View ↗


10. One stalled connection freezes the trace server ✓ Resolved
Description
main() wraps the listening socket with ctx.wrap_socket(server.socket, server_side=True), which
keeps the default do_handshake_on_connect=True. The TLS handshake therefore runs inside accept()
on the single serve_forever thread, before any worker thread starts, and the socket has no
timeout. Any peer that opens TCP to port 8080 and never finishes the handshake blocks every other
connection: probes, router reencrypt traffic and the OTel exporter. Failing probes then restart the
pod.
Code

trace_server/server.py[R979-981]

+        ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
+        ctx.load_cert_chain(TLS_CERT_FILE, TLS_KEY_FILE)
+        server.socket = ctx.wrap_socket(server.socket, server_side=True)
Relevance

●●● Strong

Blocking TLS handshakes in the listening thread create a concrete denial-of-service reliability risk
during this TLS migration.

PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
ThreadingHTTPServer only hands work to a thread after get_request() (which calls accept())
returns. On an SSLSocket created with do_handshake_on_connect=True, accept() does the
server-side handshake itself. With no timeout set, a slow or malicious client keeps the accept loop
blocked.

trace_server/server.py[973-985]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The trace server wraps its listening socket, so TLS handshakes run in the single accept loop with no timeout, and one stalled client blocks all connections.

## Fix Focus Areas
- trace_server/server.py[973-985]

## Recommended Fix
Wrap the listening socket with `do_handshake_on_connect=False`. Then do the handshake per connection inside the handler thread, e.g. override `TraceHandler.setup()` or the server's `finish_request` to call `self.request.settimeout(<n>)` and then `self.request.do_handshake()`. Also set `timeout` on the handler class so idle connections are dropped.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View medium (6)
11. Agent certificate trust is undocumented ✓ Resolved
Description
Agent deployments now mount service-ca-bundle, which the OpenShift service-ca operator populates,
but THREAT_MODEL.md has no entry describing that upstream trust dependency. When agents connect to
Valkey through the new rediss:// URL, their certificate validation relies on the operator-managed
CA without that relationship being captured in the agent threat model.
Code

openshift/deployment-backport-agent-c9s.yml[R88-90]

+      - name: service-ca
+        configMap:
+          name: service-ca-bundle
Relevance

●●● Strong

The repository accepts documenting newly introduced operational and security dependencies in
deployment documentation.

PR-#703
PR-#760

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new ConfigMap requests CA-bundle injection, and the changed agent deployment mounts it. The
changed Redis endpoint uses TLS, while the threat model contains no service-ca entry.

Rule 3951: Document new agent service dependencies in THREAT_MODEL.md
openshift/configmap-service-ca.yml[3-7]
openshift/deployment-backport-agent-c9s.yml[85-90]
openshift/configmap-endpoints-env.yml[4-7]
THREAT_MODEL.md[11-20]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Agent deployments newly rely on OpenShift service-ca for certificate trust, but the agent threat model does not describe that dependency.

## Fix Focus Areas
- THREAT_MODEL.md[11-20]
- openshift/configmap-service-ca.yml[3-7]
- openshift/deployment-backport-agent-c9s.yml[88-90]

## Recommended Fix
Add an agent service-dependency entry to THREAT_MODEL.md naming OpenShift service-ca, explaining that it supplies the CA bundle used to verify Valkey certificates, and identifying it as an upstream trust dependency for agents.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


12. Rotation runbook skips a Redis consumer ✓ Resolved
Description
Step 3 of the Annual Rotation Procedure lists the client deployments to restart. It leaves out
deployment/supervisor-processor, even though that deployment now mounts the service-ca-bundle
ConfigMap and connects to Valkey over TLS. Operators who follow the runbook (or the Jira ticket
copied from it) leave that pod on the old CA bundle after a CA rotation, so it breaks once the old
CA expires.
Code

docs/tls-compliance.md[R117-124]

+oc rollout restart deployment/api
+oc rollout restart deployment/phoenix
+oc rollout restart deployment/redis-commander
+oc rollout restart deployment/backport-agent-c9s deployment/backport-agent-c10s
+oc rollout restart deployment/rebase-agent-c9s deployment/rebase-agent-c10s
+oc rollout restart deployment/rebuild-agent-c9s deployment/rebuild-agent-c10s
+oc rollout restart deployment/mr-consolidation-agent-c9s deployment/mr-consolidation-agent-c10s
+oc rollout restart deployment/triage-agent deployment/reproducer-agent
Relevance

●●● Strong

Runbook omissions affecting CA rotation are actionable documentation defects; similar deployment
documentation updates were accepted.

PR-#703
PR-#760

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
deployment-supervisor-processor.yml gains a service-ca volume and mount in this PR and uses
REDIS_URL from endpoints-env. It does not appear in the Step 3 restart list.

openshift/deployment-supervisor-processor.yml[58-79]
docs/tls-compliance.md[113-125]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The rotation runbook's client restart list leaves out supervisor-processor, which also trusts the service CA.

## Fix Focus Areas
- docs/tls-compliance.md[117-124]

## Recommended Fix
Add `oc rollout restart deployment/supervisor-processor` to Step 3, and update the Jira ticket template in the PR description to match.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


13. The rotation check has invalid quoting ✓ Resolved
Description
The Step 4 verification loop in tls-compliance.md uses typographic quotes around its JSONPath
newline expression instead of ASCII single quotes. When operators copy the command during annual
certificate rotation, the expression may fail to parse or print stray characters, preventing clean
certificate-expiry output for the regenerated Secrets.
Code

docs/tls-compliance.md[132]

+  oc get secret "$secret" -o jsonpath="${secret}: expiry={.metadata.annotations.service\\.beta\\.openshift\\.io/expiry} version={.metadata.resourceVersion}{‘\\n’}"
Relevance

●●● Strong

Unicode shell quotes are a deterministic copy-paste correctness defect in a newly added operational
command.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Line 132 of the newly documented verification step contains the Unicode quote characters ‘ and ’
around the newline expression, and the procedure directs operators to run that command for each
regenerated Secret. The version of the command in the PR description uses ASCII single quotes
instead.

docs/tls-compliance.md[127-134]
docs/tls-compliance.md[130-134]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The annual certificate-rotation verification command uses typographic quotes in its JSONPath newline expression, so it cannot be reliably copied to check certificate expiry.

## Fix Focus Areas
- docs/tls-compliance.md[127-134]

## Recommended Fix
Replace the typographic quotes with ASCII single quotes so the expression is `{'\n'}`, using a single backslash as in the PR description. Verify the complete command in a shell.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


14. Queue-management commands stop working ✓ Resolved
Description
deployment-valkey.yml disables the plaintext port, but the queue-management commands still invoke
valkey-cli without TLS options. After this deployment, the Make targets and error-list and requeue
scripts send plaintext commands to port 6379 and cannot inspect or repair queues.
Code

openshift/deployment-valkey.yml[R22-25]

+        - --tls-port
+        - "6379"
+        - --port
+        - "0"
Relevance

●●● Strong

Existing queue-management commands must add TLS options after disabling Valkey plaintext; similar
operational script defects were accepted.

PR-#584

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new Valkey arguments set --tls-port 6379 and --port 0. Existing Make recipes and both
operational scripts execute valkey-cli against that deployment without --tls.

openshift/deployment-valkey.yml[21-31]
openshift/Makefile[16-34]
openshift/scripts/error_list.py[48-65]
openshift/scripts/requeue_error.py[54-68]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Valkey now accepts only TLS, but the repository's operational queue commands still use plaintext CLI connections.
## Fix Focus Areas
- openshift/deployment-valkey.yml[21-31]
- openshift/Makefile[16-34]
- openshift/scripts/error_list.py[48-65]
- openshift/scripts/requeue_error.py[54-68]
## Recommended Fix
Add `--tls` and the mounted service-CA certificate path to every `valkey-cli` invocation, including the documented manual commands.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


15. Missing secrets prevent failure alerts 🐞 Bug ☼ Reliability 👈 Shift left
Description
cronjob-tls-cert-check.yml mounts all three serving Secrets as required volumes, so the checker
container cannot start if any one Secret is absent. In that case check_cert() never reaches its
missing-file Sentry event, and no certificate result is reported by the Job.
Code

openshift/cronjob-tls-cert-check.yml[R69-72]

volumes:
- name: tls-phoenix-db
secret:
secretName: phoenix-db-tls # pragma: allowlist secret
Relevance

●●● Strong

Required Secret volumes can prevent the checker from starting, defeating its documented
missing-certificate alerting behavior.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Each Secret volume names a serving Secret without optional: true; the checker’s missing-file event
exists only inside the container process. A missing required Secret prevents that process from
running.

openshift/cronjob-tls-cert-check.yml[69-78]
ymir/common/tls_cert_check.py[24-37]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A missing serving Secret blocks the checker pod before its missing-certificate alert code can run.
## Fix Focus Areas
- openshift/cronjob-tls-cert-check.yml[69-78]
- ymir/common/tls_cert_check.py[24-37]
## Recommended Fix
Make each Secret volume optional so the container can start when a Secret is absent. Keep the existing file-existence check to emit a critical Sentry event for the missing certificate.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Dismiss ↗ | View ↗


16. Trace traffic fails after certificate expiry 🐞 Bug ☼ Reliability 👈 Shift left
Description
main() loads the serving certificate into its TLS context once, before serve_forever(). The
deployment mounts a service-generated certificate secret but has no reload or secret-change rollout,
so replacing the mounted files leaves a running trace server presenting the old certificate until
restart.
Code

trace_server/server.py[R979-981]

ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
ctx.load_cert_chain(TLS_CERT_FILE, TLS_KEY_FILE)
server.socket = ctx.wrap_socket(server.socket, server_side=True)
Relevance

●●● Strong

Accepted TLS reliability concern; service-ca rotates mounted certificates, but startup-only loading
leaves running servers presenting expired certificates.

PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The Service requests a generated serving secret, while the server loads its certificate only once
and the Deployment contains no certificate-reload mechanism.

openshift/service-otel-collector.yml[3-6]
openshift/deployment-otel-collector.yml[66-72]
trace_server/server.py[973-985]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The trace server keeps using its startup TLS certificate after the serving secret rotates.
## Fix Focus Areas
- trace_server/server.py[973-985]
- openshift/deployment-otel-collector.yml[66-72]
- openshift/deployment-otel-collector.yml[124-129]
## Recommended Fix
Arrange a rollout when the serving secret changes, or reload the TLS certificate and key into the server context when their mounted files change.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Dismiss ↗ | View ↗



Informational

17. A missing CA bundle fails without explanation 🐞 Bug ◔ Observability ⭐ New
Description
redis_client passes ssl_ca_certs only when os.path.isfile(ca_path) is true and otherwise
quietly falls back to the system trust store, logging nothing. If a pod is deployed without the
service-ca-bundle mount, or REDIS_TLS_CA_CERT_FILE is set wrong, every Valkey connection fails
with a generic certificate verification error. Nothing in the logs points at the missing CA file.
Code

ymir/common/base_utils.py[R64-68]

+    if redis_url.startswith("rediss://"):
+        ca_path = os.environ.get("REDIS_TLS_CA_CERT_FILE", "/etc/pki/service-ca/service-ca.crt")
+        if os.path.isfile(ca_path):
+            ssl_kwargs["ssl_ca_certs"] = ca_path
+    client = redis.Redis.from_url(redis_url, socket_timeout=socket_timeout, **ssl_kwargs)
Relevance

●●● Strong

Adding explicit diagnostics is a localized reliability improvement, consistent with accepted
error-handling findings in shared utilities.

PR-#675

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Every Valkey consumer (API, supervisor work_queue, fetcher, sweep, agents) goes through
redis_client. The isfile check has no else branch and no log line, so a missing CA file is
invisible until the TLS handshake fails.

ymir/common/base_utils.py[63-68]
ymir/api/server.py[59-62]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`redis_client` silently skips the CA bundle when the file is missing, so later TLS failures are hard to diagnose.

## Fix Focus Areas
- ymir/common/base_utils.py[63-68]

## Recommended Fix
Add an `else:` branch that calls `logger.warning("Redis TLS CA bundle %s not found; falling back to system trust store", ca_path)`. Or raise an error if `REDIS_TLS_CA_CERT_FILE` was set explicitly but the file is missing.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Dismiss ↗ | View ↗


18. Internal trace export skips certificate checks ✗ Dismissed
Description
The otlphttp/trace-server exporter now uses https://localhost:8080 with
tls.insecure_skip_verify: true, so the collector accepts any certificate the trace server
presents. Hostname verification fails anyway because the serving cert covers
otel-collector.<ns>.svc, not localhost. The setting turns off all verification instead of
pinning the service CA with an overridden server name, so this hop is encrypted but unauthenticated.
Code

openshift/configmap-otel-collector-config.yml[R18-21]

+        endpoint: https://localhost:8080
        encoding: json
+        tls:
+          insecure_skip_verify: true
Relevance

●●● Strong

TLS compliance intent favors fixing unauthenticated internal trace traffic; recent security findings
were accepted.

PR-#806

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The serving cert comes from the otel-collector Service annotation, so it is issued for the service
DNS name and not for localhost. The PR therefore turned verification off instead of configuring a CA
and a server-name override.

openshift/configmap-otel-collector-config.yml[15-21]
openshift/service-otel-collector.yml[1-6]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The OTel exporter to the trace-server sidecar disables TLS verification completely.

## Fix Focus Areas
- openshift/configmap-otel-collector-config.yml[17-21]
- openshift/deployment-otel-collector.yml[106-129]

## Recommended Fix
Mount `service-ca-bundle` into the otel-collector container. Replace `insecure_skip_verify: true` with `ca_file: /etc/pki/service-ca/service-ca.crt` and `server_name_override: otel-collector.jotnar-ymir--jotnar-ymir.svc`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
✅ Compliance rules (platform): 5 rules
Review mode: Auto: 🧠 Deep: Substantial logic changes span many independent workflows and deployment paths.

Grey Divider

Tip of the day
💡 Did you know, you can keep summaries lean with Findings visible per group, which tucks the rest behind a View link

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread openshift/deployment-backport-agent-c9s.yml
Comment thread openshift/deployment-valkey.yml
Comment thread docs/tls-compliance.md Outdated
Comment thread openshift/configmap-endpoints-env.yml Outdated
Comment thread openshift/deployment-phoenix.yml Outdated
Comment thread trace_server/server.py Outdated
Comment thread docs/tls-compliance.md
@majamassarini
majamassarini force-pushed the tls-compliance-remediation branch from 56f6ff0 to df40e31 Compare October 8, 2026 13:35
Enable HTTP Strict Transport Security on all externally exposed Routes
to enforce HTTPS in browsers for 1 year, addressing a compliance gap
for the Red Hat encryption-in-transit requirement.

Assisted-by: Claude Haiku 4.5 <noreply@anthropic.com>
Document the encryption-in-transit posture of the Ymir deployment
covering external Routes (HSTS, TLS termination), database/cache
TLS, SDN isolation for internal HTTP services, and PFS cipher
compliance.

Assisted-by: Claude Haiku 4.5 <noreply@anthropic.com>
Configure the phoenix-db deployment to serve TLS using certificates
auto-generated by the OpenShift service-ca operator. An initContainer
copies the cert/key to a writable volume with correct permissions and
generates a PostgreSQL SSL config snippet. The Phoenix application
now connects with sslmode=verify-full to ensure encrypted and
authenticated database connections.

Assisted-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Configure Valkey to serve TLS-only connections using certificates
auto-generated by the OpenShift service-ca operator. Update REDIS_URL
to use the rediss:// scheme and configure redis-commander for TLS.

Assisted-by: Claude Sonnet 4.5 <noreply@anthropic.com>
Add TLS support to the trace-server Python application via environment-
configurable cert/key paths. Mount OpenShift service-ca certificates in
the deployment and switch both trace-server Routes from edge to reencrypt
termination so router-to-pod traffic is encrypted. Update OTel collector
config to export via HTTPS to the now-TLS trace-server sidecar.

Assisted-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add service-ca-bundle ConfigMap volume mount to all deployments
- Valkey mTLS: add --tls-auth-clients no (clients have no certs)
- Python Redis client: pass ssl_ca_certs from mounted CA bundle
- phoenix-db init container: add ImageStream trigger for image resolution
- Update all CA cert paths to /etc/pki/service-ca/service-ca.crt

Assisted-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Service CA is valid 26 months and auto-rotates at 13 months with a
13-month grace period. An annual proactive rotation (delete secrets,
restart pods) is simpler and more reliable than monitoring expiry
thresholds. Replace the CronJob-based checker with a documented
runbook tracked by a recurring Jira issue.

Assisted-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Use FQDN (*.jotnar-ymir--jotnar-ymir.svc) in REDIS_URL,
  REDIS_HOST, and PHOENIX_SQL_DATABASE_URL so hostnames match the
  service-ca certificate SANs (redis-py 6.x checks by default)
- Add TLS flags to all valkey-cli invocations in Makefile and
  requeue_error.py (Valkey is now TLS-only, no plaintext port)
- Fix trace server TLS: defer handshake to worker threads so a
  stalled client cannot block all connections and health probes
- Fix typographic curly quotes in docs JSONPath command
- Add supervisor-processor and mcp-gateway to rotation runbook
- Document service-ca trust dependency in THREAT_MODEL.md
- Add TLS flags to error_list.py direct valkey-cli invocation

Assisted-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@majamassarini
majamassarini force-pushed the tls-compliance-remediation branch from df40e31 to 5387283 Compare October 8, 2026 13:37
@majamassarini

Copy link
Copy Markdown
Member Author

/agentic_review

Comment thread openshift/deployment-redis-commander.yml
Comment thread openshift/configmap-otel-collector-config.yml
Comment thread ymir/common/base_utils.py
Comment on lines +64 to +68
if redis_url.startswith("rediss://"):
ca_path = os.environ.get("REDIS_TLS_CA_CERT_FILE", "/etc/pki/service-ca/service-ca.crt")
if os.path.isfile(ca_path):
ssl_kwargs["ssl_ca_certs"] = ca_path
client = redis.Redis.from_url(redis_url, socket_timeout=socket_timeout, **ssl_kwargs)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Informational

18. A missing ca bundle fails without explanation 🐞 Bug ◔ Observability

redis_client passes ssl_ca_certs only when os.path.isfile(ca_path) is true and otherwise
quietly falls back to the system trust store, logging nothing. If a pod is deployed without the
service-ca-bundle mount, or REDIS_TLS_CA_CERT_FILE is set wrong, every Valkey connection fails
with a generic certificate verification error. Nothing in the logs points at the missing CA file.
Agent Prompt
## Issue description
`redis_client` silently skips the CA bundle when the file is missing, so later TLS failures are hard to diagnose.

## Fix Focus Areas
- ymir/common/base_utils.py[63-68]

## Recommended Fix
Add an `else:` branch that calls `logger.warning("Redis TLS CA bundle %s not found; falling back to system trust store", ca_path)`. Or raise an error if `REDIS_TLS_CA_CERT_FILE` was set explicitly but the file is missing.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Dismiss ↗ | View ↗

@qodo-for-packit

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 5387283

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant