An autonomous Kubernetes controller that uses a local LLM to analyze real-time cluster metrics and make intelligent infrastructure decisions. Instead of blindly restarting failing pods, CourtVision reasons about resource contention, noisy neighbor problems, and capacity constraints β then recommends whether to adjust resource limits, migrate pods to different nodes, or scale deployments.
CourtVision runs in two shapes. Single-cluster mode watches one cluster with one agent. Multi-agent mode runs one subagent per cluster, each monitoring its own cluster in parallel, plus a master coordinator that reasons across the whole fleet β spotting cross-cluster opportunities like relieving an overloaded cluster by shifting work to one with spare capacity.
CourtVision runs a continuous monitoring loop that collects resource metrics from your Kubernetes cluster every few seconds, feeds them to a local LLM (Llama 3 via Ollama), and surfaces structured decisions through a REST API and real-time dashboard.
βββββββββββββββ ββββββββββββββββ βββββββββββββ βββββββββββββββββ
β Kubernetes ββββββΆβ CourtVision ββββββΆβ Ollama ββββββΆβ Decisions β
β Cluster β β Agent β β (LLM) β β + Dashboard β
β βββββββ βββββββ β β β
β metrics- β β - collect β β Llama 3 β β - REST API β
β server β β - analyze β β local β β - SSE stream β
β β β - decide β β β β - React UI β
βββββββββββββββ ββββββββββββββββ βββββββββββββ βββββββββββββββββ
The LLM doesn't just detect problems β it explains its reasoning in natural language and chooses the optimal remediation action from a set of available operations:
- patch_limits β adjust CPU/memory limits when a pod is near capacity but the node has headroom
- evict_and_move β migrate a pod to a less-loaded node when the current node is under pressure
- scale_down β reduce replicas when a deployment is over-provisioned
- none β continue monitoring when metrics are elevated but not dangerous
For fleets with more than one cluster, courtvision multi-monitor runs a two-tier topology:
ββββββββββββββββββββββββββββ
β Coordinator (master) β slow loop (~30s)
β reads cached snapshots β cross-cluster reasoning
β β fleet-wide decisions β β /api/decisions
ββββββββββββββ¬ββββββββββββββ
reads LatestSnapshot() β (never touches clusters directly)
ββββββββββββββββββββββββββΌβββββββββββββββββββββββββ
βΌ βΌ βΌ
βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ
β ClusterWorker β β ClusterWorker β β ClusterWorker β fast loop (~5s)
β prod-us β β prod-eu β β staging β collectβanalyzeβstore
β own store+exec β β own store+exec β β own store+exec β /api/clusters/{name}/β¦
ββββββββββ¬βββββββββ ββββββββββ¬βββββββββ ββββββββββ¬βββββββββ
βΌ βΌ βΌ
cluster A cluster B cluster C
- Subagents (
ClusterWorker) β one per cluster, each running the full collect β analyze β store pipeline on its own fast loop. Each owns its own decision store and an executor bound to that cluster's kubeconfig context, so any auto-safe action lands on the right cluster. - Master agent (
Coordinator) β runs a deliberately slower loop. It reads each worker's cached snapshot (it never collects from clusters itself), asks the LLM to reason about the fleet as a whole, and records cross-cluster decisions in its own store. It only runs once at least two clusters have reported, so cold start and single-cluster cases stay quiet. - Cluster attribution β coordinator decisions carry a
target_clusterso the fleet view can attribute each cross-cluster recommendation to the cluster it targets. These decisions are surfaced read-only (advisory); the HTTP API never executes them β auto-safe is per-cluster and reversible-only, and cross-cluster moves stay operator-driven.
The two loops are independently tunable (--interval for workers, --coordinator-interval for the master). The master is meant to run several times slower than the workers β it reads cached state, so running it faster buys nothing, and cross-cluster moves are strategic decisions that shouldn't be re-litigated every few seconds.
- Real Kubernetes integration β connects to any cluster via kubeconfig (AWS EKS, GKE, AKS, Minikube, Kind)
- Multi-cluster, multi-agent β one subagent per cluster plus a master coordinator for cross-cluster reasoning, all in a single process
- Local LLM analysis β uses Ollama with Llama 3 for on-device inference, no data leaves your machine
- Interactive CLI β styled REPL with a
/command palette (filter, arrow-navigate, Tab-complete) and switchable session modes, plus colored output and spinners - Real-time dashboard β React frontend with glassmorphism UI, live metric visualization, and SSE-powered decision feed
- Mock mode β full demo experience without a cluster or LLM, using simulated metrics with a noisy neighbor scenario
- Dry-run by default β decisions are proposed and displayed but never executed unless explicitly enabled
- Rule-based fallback β deterministic engine catches critical issues even if the LLM is unavailable
git clone https://github.com/atulya-singh/CourtVision.git
cd CourtVision
go build -o courtvision ./cmd/courtvision/go install github.com/atulya-singh/CourtVision/cmd/courtvision@latest- Ollama β install from ollama.com, then run
ollama pull llama3 - Kubernetes cluster (optional) β any cluster with metrics-server installed. For local testing, use Kind:
kind create cluster
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
kubectl patch deployment metrics-server -n kube-system --type='json' \
-p='[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'courtvisionThis drops you into the CourtVision REPL. Press / to open a command
palette β a live-filtered, arrow-navigable menu you drive with
almost no typing β and switch session modes once instead of re-typing flags:
β CourtVision v1.0.0
β ollama Β· metrics:mock Β· ns:all Β· sandbox
βΊ / β press "/" to open the palette
βΈ /analyze Run a one-shot cluster analysis
/review Analyze, then approve/reject each action inline
/metrics <mock|k8s> Switch the metrics source for this session
/namespace <ns|all> Set the namespace filter for this session
β¦ (β/β move Β· Tab complete Β· Enter run Β· Esc dismiss)
- Palette β type
/then filter (/anβ/analyze);β/βhighlight,Tabcompletes,Enterruns,Escdismisses.β/βstill cycle command history when the palette is closed. - Modes β
/metrics k8s,/namespace kube-system, orShift+Tabto toggle mockβk8s. The mode bar above the prompt always shows the current context, and/review//analyzerun under it (no repeated--metrics/--namespace). - Back-compat β bare words (
review,status,help,exit) still work. Long-running servers (/monitor,/multi-monitor) print the shell command to run rather than taking over the REPL. The REPLreviewflow is a sandbox (k8s falls back to dry-run), so switching/metrics k8snever mutates a cluster.
βΊ /review
Analyzing cluster...
Review 1/2 [critical]
Pod: default/data-pipeline-545bf66bc6
Action: patch_limits
Why: Pod consuming 100% of CPU limit, recommend raising the ceiling
[a]pprove [r]eject [s]kip [A]pprove all [tab] auto: off [q]uit
βΊ exit
# Check that Ollama and Kubernetes are reachable
courtvision status
# Quick cluster analysis β print results and exit
courtvision analyze --metrics k8s --namespace production --output table
# Same analysis as JSON (for piping to jq or other tools)
courtvision analyze --metrics k8s --output json
# Start continuous monitoring with dashboard
courtvision monitor --metrics k8s --namespace production --port 8080
# Run with mock data (no cluster needed)
courtvision monitor --metrics mock --port 8080
# Multi-agent mode β one subagent per cluster + a cross-cluster coordinator
courtvision multi-monitor --clusters prod-us,prod-eu,staging --metrics k8s --interval 5s
# Multi-agent mode with mock data (two simulated clusters, no cluster or LLM needed)
courtvision multi-monitor --clusters mock-a,mock-b --metrics mockIn multi-agent mode the API exposes per-cluster state and the coordinator's fleet-wide decisions:
curl localhost:8080/api/clusters # fleet roll-up (pods/nodes/decisions per cluster)
curl localhost:8080/api/clusters/prod-us/snapshot # one cluster's latest snapshot
curl localhost:8080/api/clusters/prod-us/decisions # one cluster's decisions
curl localhost:8080/api/decisions # coordinator's cross-cluster decisionsWhen the monitor is running, open http://localhost:8080 for the API, or start the React dashboard:
cd web
npm install
npm run devThen open http://localhost:5173 to see the real-time dashboard with cluster visualization and streaming LLM decisions.
Start the continuous monitoring agent with a read-only API server and
dashboard. The dashboard observes metrics and decisions but never mutates the
cluster; to execute a decision use analyze --apply, the REPL review flow, or
multi-monitor --auto-safe.
| Flag | Default | Description |
|---|---|---|
--metrics |
mock |
Metrics source: mock or k8s |
--namespace |
`` (all) | Kubernetes namespace to watch |
--port |
8080 |
API server port |
--ollama-url |
http://localhost:11434 |
Ollama server URL |
--model |
llama3 |
LLM model name |
--interval |
3s |
Monitoring loop interval |
Monitor multiple clusters with one subagent each plus a cross-cluster coordinator.
| Flag | Default | Description |
|---|---|---|
--clusters |
(required) | Comma-separated kubeconfig context names to monitor |
--metrics |
mock |
Metrics source: mock or k8s |
--namespace |
`` (all) | Kubernetes namespace to watch (applies to every cluster) |
--port |
8080 |
API server port |
--ollama-url |
http://localhost:11434 |
Ollama server URL |
--model |
llama3 |
LLM model name |
--interval |
5s |
Per-cluster worker loop interval |
--coordinator-interval |
30s |
Coordinator (cross-cluster) loop interval |
--dry-run |
true |
Log decisions without executing |
--auto-safe |
false |
Let each worker auto-execute its own reversible decisions (cordon_node, scale_down, patch_limits); evict_and_move still waits for approval |
--auto-cooldown |
3m |
In auto-safe mode, suppress repeat auto-execution of the same action on the same target for this long |
--audit-log |
`` (off) | Append a durable JSONL record of every executed action (all clusters) to this file |
--audit-max-bytes |
0 (unbounded) |
Rotate the audit log past this size, keeping a few numbered backups |
The audit trail is also served read-only at
GET /api/audit(fleet) andGET /api/clusters/{cluster}/audit(per-cluster), newest-first β live even without--audit-log(backed by an in-memory ring).
Auto-safe is the unattended tier: each worker heals its own cluster's reversible issues without a human in the loop, while the coordinator's cross-cluster moves stay advisory (surfaced read-only, no auto-execution).
--dry-runis still the separate safety switch β--auto-safewith--dry-run=falseperforms real autonomous mutations on live clusters (the startup banner warns when both are set). The per-target cooldown stops the ~5s analysis loop from re-firing the same fix every tick.
Run a one-shot cluster analysis and exit.
| Flag | Default | Description |
|---|---|---|
--metrics |
mock |
Metrics source: mock or k8s |
--namespace |
`` (all) | Kubernetes namespace to watch |
--output |
table |
Output format: table or json |
--ollama-url |
http://localhost:11434 |
Ollama server URL |
--model |
llama3 |
LLM model name |
Verify the tamper-evidence hash chain of a JSONL audit log written with --audit-log. Reports the first event whose content was altered or whose link was broken (edited, reordered, or removed), or confirms the chain is intact. Exits non-zero on tampering, so it drops into a cron/CI check. Verify a single unrotated file β each rotated segment (<file>.1, .2, β¦) carries its own independent chain.
Check connectivity to Ollama and Kubernetes.
Print version, commit hash, and build date.
CourtVision/
βββ cmd/courtvision/ β CLI entry point (Cobra commands)
β βββ main.go β root command + REPL
β βββ monitor.go β continuous (single-cluster) monitoring subcommand
β βββ multi.go β multi-cluster monitoring subcommand (workers + coordinator)
β βββ analyze.go β one-shot analysis subcommand
β βββ status.go β connectivity check subcommand
βββ internal/
β βββ types/ β shared data structures (PodMetrics, Decision, etc.)
β βββ metrics/
β β βββ mock.go β simulated cluster with noisy neighbor
β β βββ k8s.go β real Kubernetes metrics via client-go (kubeconfig context aware)
β βββ llm/
β β βββ client.go β Ollama HTTP client
β β βββ engine.go β LLM decision engine (implements Engine interface)
β β βββ prompt.go β cluster snapshot β structured LLM prompt (+ cross-cluster prompt)
β β βββ parser.go β LLM text output β structured decisions
β βββ decision/
β β βββ engine.go β rule-based fallback engine
β βββ cluster/
β β βββ worker.go β ClusterWorker: per-cluster subagent pipeline
β β βββ coordinator.go β Coordinator: master agent for cross-cluster reasoning
β βββ store/
β β βββ store.go β thread-safe shared state with SSE pub/sub
β βββ api/
β β βββ server.go β REST API + SSE endpoints (single- and multi-cluster routes)
β βββ ui/
β βββ styles.go β terminal styling (lipgloss colors, layouts)
βββ web/ β React dashboard (Vite + TypeScript + Tailwind)
- Metrics Provider (
mock.goork8s.go) collects a cluster snapshot every N seconds, stamped with the cluster's name - LLM Engine converts the snapshot to a prompt, sends it to Ollama, parses the response into structured decisions
- Store saves the snapshot and decisions, notifies SSE subscribers
- API Server serves cluster state via REST and streams decisions via SSE β read-only; it never mutates the cluster
- Dashboard (React) renders the cluster visually and shows the decision feed in real-time
In multi-agent mode this pipeline runs once per cluster inside a ClusterWorker, each with its own store and executor. A Coordinator sits on top: it periodically reads every worker's cached snapshot, builds a single cross-cluster prompt, and records fleet-wide decisions in its own store. Because cluster identity flows through ClusterSnapshot and Decision, decisions are always attributable to the cluster they belong to; execution happens only through the CLI or a worker's own auto-safe loop, never through the API.
Every major component is behind an interface β metrics.Provider, decision.Engine, llm.Generatable. This means you can swap implementations without changing any other code. Mock metrics β real Kubernetes. Rule engine β LLM engine. Local Ollama β remote API. One line change in the wiring, zero changes elsewhere.
This is exactly what makes multi-agent mode cheap: a ClusterWorker is just the single-cluster pipeline behind those same interfaces, parameterized by a kubeconfig context. Spinning up N clusters is N instances of the same wiring, and the Coordinator composes them without any of them knowing it exists.
- Go β agent, CLI, API server
- client-go β Kubernetes API client
- Ollama + Llama 3 β local LLM inference
- Cobra β CLI framework
- Lipgloss / Bubbletea β terminal UI styling
- React + TypeScript + Tailwind β dashboard frontend
- Server-Sent Events β real-time streaming