AI inference and tooling running on a three-node SNUC cluster under OpenShift AI.
This is an edge cluster of AMD Ryzen AI 9 HX370 PCs, and that shapes almost everything in here. There are no discrete GPUs — just one Radeon 890M integrated GPU per node, sharing system RAM. Which accelerator backend works, how much memory a pod may request, and which models can be served at all are decided by that hardware, not by preference. Those constraints are documented in Hardware Notes; read them before adding a workload.
OpenShift 4.21 with OpenShift AI 3.5, three nodes, RHEL CoreOS 9.6.
| snuc-01 | snuc-02 | snuc-03 | |
|---|---|---|---|
| CPU | 24 vCPU | 24 vCPU | 24 vCPU |
| System RAM | 62 GB allocatable | 62 GB | 46 GB |
| GPU | Radeon 890M | Radeon 890M | Radeon 890M |
| Dedicated VRAM | 32 GB | 32 GB | 48 GB |
Every node is an AMD Ryzen AI 9 HX 370 with a Radeon 890M iGPU (gfx1150). One GPU each, so three GPUs total — which is the binding constraint on how many accelerated workloads can run at once.
| Component | Description |
|---|---|
| llm/llamacpp | llama.cpp as an OpenAI-compatible KServe inference server, Vulkan GPU acceleration. |
| llm/lemonade | Lemonade as an alternative GGUF backend, also Vulkan. |
| litellm | LiteLLM proxy + Postgres, for routing to external model providers. |
| command-center | Web UI for managing and chatting with multiple AI models from one interface. |
| vision-ai | Real-time object detection and instance segmentation (YOLO26) with a streaming web app. |
Command Center (pictured above) is the front door to all of this: KServe InferenceServices and external providers appear side by side in one model list, with per-response latency stats (TTFT, token count, prompt tokens) and guardrail enforcement on blocked content.
┌──────────────────────────────────────────┐
│ command-center │
│ FastAPI + static UI, model proxy │
└───────┬─────────────────────┬────────────┘
│ │
KServe ISVc│ │OpenAI-compatible
discovery │ │ (external models)
▼ ▼
┌───────────────────────────────────┐ ┌──────────────┐
│ KServe InferenceServices │ │ litellm │
│ │ │ proxy │
│ llamacpp │ lemonade │ YOLO26 │ └──────────────┘
│ (GGUF) │ (GGUF) │ (ONNX) │
└───────────────────────────────────┘
▲
gRPC │ :8001
│
┌──────────────┐ ┌───────┴────────┐
│ RTSP camera │─────▶│ vision-ai app │─────▶ annotated MJPEG stream,
│ / video file │ │ OpenCV+FastAPI │ snapshots, JSON API
└──────────────┘ └────────────────┘
Each component owns its own k8s/ manifests, pipelines/, and README. The component
README is authoritative for build and deployment steps.
Two interchangeable backends serve GGUF models through KServe with an OpenAI-compatible
API on port 8080. Both auto-discover a .gguf file from the KServe model mount path.
| llamacpp | lemonade | |
|---|---|---|
| Server | llama-server |
lemond |
| Acceleration | Vulkan | Vulkan |
| Model alias | filename | registered under the InferenceService name |
Tunables (CONTEXT_LENGTH, N_GPU_LAYERS, REASONING_BUDGET, FLASH_ATTENTION,
EXTRA_ARGS) are set as ServingRuntime defaults and overridable per-model on the
InferenceService. See llm/llamacpp.
YOLO26 models served via KServe, consumed over gRPC by the vision-ai app.
| Model | Purpose | Runtime |
|---|---|---|
| YOLO26m | Object detection (COCO 80 classes) | onnxruntime-cpu |
| YOLO26n-seg | Detection + instance segmentation | onnxruntime-cpu |
An onnxruntime-migraphx ServingRuntime also exists for GPU inference, but is not
currently in use — see the hardware note below.
The 890M has no memory of its own. Dedicated VRAM is carved out of system RAM in BIOS, so raising it directly reduces what pods can request on that node. snuc-03's 48 GB carve-out costs it 16 GB of schedulable RAM (46 GB vs 62 GB elsewhere) — enough that a 4 GiB pod request no longer fits there.
Current GTT is 30.9 GB on snuc-01/02 and 23.0 GB on snuc-03. Read these per node with:
cat /sys/class/drm/card0/device/mem_info_vram_total # dedicated VRAM
cat /sys/class/drm/card0/device/mem_info_gtt_total # GTT (shared)Since ROCm does not work here anyway (below), a larger VRAM reservation currently buys nothing and only costs schedulable RAM.
ONNX Runtime with ROCMExecutionProvider fails at startup with
Hip error: 'out of memory' in hipblaslt init. This is not a VRAM sizing problem.
It was tested directly on both a 32 GB node and the 48 GB node, at container memory
limits from 8 GiB to 12 GiB, and fails identically every time. The AMD GPU is
successfully allocated to the pod — the failure is inside hipblaslt on gfx1150.
Consequences:
- GGUF backends (llamacpp, lemonade) use Vulkan, which addresses shared system memory and works fine.
- ONNX vision models run on
onnxruntime-cpu. Theonnxruntime-migraphxServingRuntime exists but is unused; note its image shipsROCMExecutionProviderandCPUExecutionProviderbut no MIGraphX provider, so naming MIGraphX in--providerssilently falls through to CPU while still holding a GPU. - Pair CPU-runtime models with the
cpu-onlyhardware profile so they do not reserve a GPU they cannot use. That profile is defined in vision-ai/hardware-profiles/, not shipped by OpenShift AI.
llm/
llamacpp/ # llama.cpp serving runtime + KServe manifests
lemonade/ # Lemonade serving runtime + KServe manifests
litellm/ # LiteLLM proxy manifests
command-center/ # Web UI (app, k8s/, pipelines/)
vision-ai/ # Vision inference (app/, k8s/, pipelines/, argocd/, hardware-profiles/)
audio-samples/ # Shared test fixtures
video-samples/
Tekton pipelines build each component and restart its deployment; ArgoCD syncs the
k8s/ manifests. A shared github-listener EventListener in the visionai namespace
routes pushes to the right pipeline by branch.
- OpenShift with OpenShift AI 3.5 installed
- OpenShift GitOps (ArgoCD) and OpenShift Pipelines (Tekton) operators
- S3-compatible object storage (e.g., OpenShift Data Foundation) for model artifacts
- A container registry you can push to (this repo assumes an in-cluster Quay)
- An accelerator, if you want one that works. On discrete AMD or NVIDIA GPUs the
ROCm limitation above does not apply, and the
onnxruntime-migraphxruntime andamd-gpu-visionhardware profile become usable as written.
These are baked into the manifests and will not match your cluster. Grep and replace before applying anything:
| What | Current value |
|---|---|
| Workload namespace | john |
| Pipeline namespace | visionai |
| Registry host | registry-quay-app.quay.svc.cluster.local (in-cluster) and quay.apps.snuc.kubernetes.day (route) |
| Cluster domain | apps.snuc.kubernetes.day |
| Git remote | github.com/explicitworkload/snuc-openshift-ai |
Credentials are committed in plaintext in
litellm/litellm-deployment.yaml(LITELLM_MASTER_KEY, the Postgres DSN) andcommand-center/k8s/deployment.yaml(COMMAND_PASSWORD). These are demo values. Replace them — ideally move them to Secrets — before exposing anything.
oc new-project john # workloads: models, apps
oc new-project visionai # Tekton pipelines and triggersCreate an OpenShift AI data connection in john pointing at your S3 bucket. It must
produce a secret named models-odf-s3. Then create the service account the
InferenceServices use to pull weights:
oc create sa models-odf-s3-sa -n johnUpload the ONNX models to the bucket:
/models/yolo26m.onnx # detection
/models/yolo26n-seg.onnx # segmentation
GGUF models for the LLM runtimes go in the same bucket; see llm/llamacpp for the paths its InferenceService expects.
# Registry push credentials for Tekton
oc create secret docker-registry quay-push-secret \
--docker-server=<your-registry> \
--docker-username=<user> --docker-password=<pass> -n visionai
oc annotate secret quay-push-secret -n visionai \
tekton.dev/docker-0=https://<your-registry>
# GitHub webhook shared secret (only if you want push-triggered builds)
oc create secret generic github-webhook-secret \
--from-literal=secret=<random-string> -n visionai
# RTSP camera credentials (only if vision-ai reads a live camera)
oc create secret generic rtsp-credentials \
--from-literal=RTSP_URL='rtsp://<user>:<pass>@<camera-ip>/stream1' -n johnThe ServingRuntimes reference images that must exist before the models will start.
Build the vision serving images from vision-ai/ (build context is ./vision-ai):
podman build -f vision-ai/Dockerfile.rocm-base -t <registry>/visionai:rocm-base vision-ai/
podman build -f vision-ai/Dockerfile -t <registry>/visionai:latest vision-ai/
podman build -f vision-ai/Dockerfile.cpu -t <registry>/visionai:cpu vision-ai/
podman build -f vision-ai/app/Dockerfile -t <registry>/visionai-app:yolox vision-ai/
podman push <registry>/visionai:rocm-base && podman push <registry>/visionai:latest
podman push <registry>/visionai:cpu && podman push <registry>/visionai-app:yoloxDockerfile builds FROM the rocm-base tag, so push that one first. The LLM and
command-center images build the same way from their own component directories.
Models on a CPU runtime should not reserve a GPU. OpenShift AI's default-profile
declares amd.com/gpu with minCount: 1, so anything referencing it takes a GPU
whether or not it can use one. default-profile is the only profile OpenShift AI
ships, so both profiles below have to be created — the InferenceServices reference
cpu-only by name and will not start without it:
oc apply -f vision-ai/hardware-profiles/cpu-only.yaml # required
oc apply -f vision-ai/hardware-profiles/amd-gpu-vision.yaml # only if you have a usable GPUIf a GPU request was already injected into a stored InferenceService, changing the profile will not remove it —
oc applymerges maps and leaves the key in place. Strip it explicitly:oc patch inferenceservice <name> -n john --type=json -p \ '[{"op":"remove","path":"/spec/predictor/model/resources/requests/amd.com~1gpu"}, {"op":"remove","path":"/spec/predictor/model/resources/limits/amd.com~1gpu"}]'
ArgoCD owns the vision stack. Point it at your fork first, then apply:
oc apply -f vision-ai/argocd/application.yamlThe rest apply directly:
oc apply -f litellm/ # LiteLLM proxy + Postgres
oc apply -f command-center/k8s/ # Web UI (RBAC, Deployment, Service, Route)
oc apply -f llm/llamacpp/servingruntime.yaml -f llm/llamacpp/inferenceservice.yamloc apply -f vision-ai/pipelines/workspace-pvc.yaml
oc apply -f vision-ai/pipelines/pipeline.yaml
oc apply -f command-center/pipelines/pipeline.yaml
oc apply -f vision-ai/pipelines/triggers.yaml # EventListener + webhook RouteThen add the github-webhook route URL as a push webhook in your GitHub repo, using
the shared secret from step 3.
vision-ai/pipelines/triggers.yamldeclares both triggers on a shared EventListener. Removing either one from that file deletes it from the cluster on the next apply.
oc get inferenceservice -n john # all should reach READY=True
oc get pods -n john
oc get route command-center -n john # open the UIIf an InferenceService stays Progressing, check whether its pod is unschedulable —
on this cluster that is usually GPU or memory requests, not a model problem.
