Skip to content

Operator: scope secret access to the app-N namespaces it creates #299

Description

@v0l

Follow-up from #297 part 2 (PR pending). The verb trim there removes events, removes watch everywhere and cuts each rule to what the reconcile loop issues, but secret access is still cluster-wide: the ClusterRole grants secrets: get, create, patch in every namespace.

The operator only ever touches secrets in the app-N namespaces it created itself:

  • reads generated (lnvps_operator/src/app_deployments.rs:1035-1036)
  • applies generated and the per-service file secret (:1315, :1341)

A stolen operator token therefore reads every tenant's generated database passwords and every app-tls private key, plus anything in kube-system and cert-manager's namespace.

Why the obvious fix does not work

Moving secrets to a Role in each app-N namespace requires the operator to create that Role and RoleBinding at reconcile time, and Kubernetes RBAC refuses to let a subject grant permissions it does not itself hold (privilege escalation prevention). Holding them cluster-wide to be allowed to grant them namespace-wide is exactly the state we are trying to leave.

What would work

  1. A second, static ClusterRole — say lnvps-operator-appns — carrying the namespaced verbs (secrets, configmaps, services, persistentvolumeclaims, deployments, networkpolicies, ingresses, pods).
  2. lnvps-operator keeps only the cluster-scoped verbs (namespaces) plus rolebindings: create, get, patch, delete and bind on lnvps-operator-appns by resourceNames. bind is the escape hatch RBAC provides for exactly this: it permits creating a binding to a named role without holding its permissions.
  3. The operator creates a RoleBinding in each app-N namespace next to the NetworkPolicy it already applies.

Notes for whoever picks this up:

  • Rollout is two-step and ordered: both ClusterRoles must exist before the new operator image runs, or the first reconcile 403s on secrets. It is self-healing — the next pass succeeds once the binding exists — but the first pass after a fresh namespace may log a 403 while the authorizer cache catches up.
  • This narrows the blast radius of the token, not of the pod: the operator still holds the field encryption key and the database DSN, so a compromise still reads tenant config from the database. Operator hardening: non-root runtime, scoped RBAC, dedicated DB credentials #297 part 3 is the other half.
  • Needs a cluster apply from Kieran; merging alone changes nothing.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions