Platform engineer. I build agentic SDLC pipelines and the infrastructure they run on.
Thesis — bounded autonomy. Let agents do the work, but enforce in code what a prompt can only ask, ground every verdict in executed evidence, and keep a human at the irreversible step. Every repo below is that idea under a different constraint.
autonomous-dev-loop |
Issue → AI PR → evidence-backed AI review → auto-fix → human merge, on GitHub Actions. Built after an agent replaced a 26-test suite with an 18-line stub; 11 of its own features were shipped by the loop. 800+ tests, 25 ADRs. |
ai-dev-tools |
The interactive counterpart in Claude Code: typed builders, an adversarial challenger, and a reviewer that runs your tests before it calls a patch done. |
agent-patterns |
15 multi-agent patterns indexed by the problem they solve — diagrams, trade-offs, failure modes, runnable code. 12 extracted from the two repos above. |
gitops-homelab |
Chat, code, vision, speech-to-text and text-to-speech models on one 6 GB GTX 1060 — bare-metal microk8s, ArgoCD, GPU time-slicing. 25 ADRs, 13 incidents written up with root cause. |
ollama-chat |
Self-hosted chat UI for local LLMs, deployed by the repo above. |
sync-engine |
Webhooks drop events, polling is late: a Python sync engine built around 11 integration failure modes, each mapped to an ADR and a test. |
fullstack-pilot |
Node, Python and .NET on MongoDB, PostgreSQL and SQL Server — with CI that tests database image upgrades against existing data. |
flowchart LR
subgraph sdlc["Agentic SDLC"]
adl["autonomous-dev-loop<br/>headless · GitHub Actions"]
adt["ai-dev-tools<br/>interactive · Claude Code"]
ap["agent-patterns"]
end
subgraph infra["Self-hosted infrastructure"]
gh["gitops-homelab<br/>microk8s · ArgoCD · GTX 1060"]
oc["ollama-chat"]
end
subgraph fund["Platform fundamentals"]
se["sync-engine"]
fp["fullstack-pilot"]
end
adl -- "patterns extracted to" --> ap
adt -- "patterns extracted to" --> ap
adl -.-|"same pipeline, two modes"| adt
gh -- "deploys" --> oc
- Decisions in writing. 65+ ADRs across these repos, each with rejected alternatives and consequences.
- Incidents, not hypotheticals. Guardrails and config trace back to a specific failure: a deleted test suite, a CUDA OOM, a Helm setting that never reached the pod.
- Measure, then decide. Quantization chosen by benchmark and by executing the generated code, not by reputation.
- Honest scope. What's enforced in code and what's only asked of a model are documented separately.
- Moving LLM guardrails from prompts into deterministic checks — import allowlists and type-check gates before a PR opens (ADR-0019, proposed).
- Evaluating local models on constrained hardware by running what they produce (ADR-0011).
Node.js · .NET · Python · Kubernetes · ArgoCD · GitHub Actions · Claude API · Ollama



