Systems engineer working on local LLM inference: measuring how models actually behave on specific hardware, and fixing it upstream when they don't behave. My lab is an RX 9070 XT (RDNA4) desktop plus six Tesla P100s in two home servers.
I use AI to make AI, break AI, and make it again. Every day. The agents (Claude, Gemini) write the code; the hardware, the experiments, the review and the final calls are mine. Every claim below links to its evidence.
- KV caches that survive a restart. llama-server restored a saved session and then discarded it on the next request. A 117-line checkpoint sidecar fixed it: a parked 100K-token session comes back in about 4 s instead of a 12-minute re-prefill. Merged in llama-cpp-turboquant #206. Writeup
- Tesla P100s were doing quality-sensitive math in fp16 for years. A three-line fix moved median KLD against an fp32 reference from 0.0023 to 0.000001, with no speed cost. Merged in llama-cpp-turboquant #212 and buun-llama-cpp #80, and reported upstream as ggml-org/llama.cpp #25593. Writeup
- Tool calling restored for a Qwen3.8 fine-tune. The chat template it shipped with silently dropped
tool_callsfrom the conversation history. The model's author adopted the fixed template (discussion). Template and test report - CUDA-on-AMD defects on RDNA4. Reported two defects in SCALE on gfx1201:
- Other merged work:
- AMD Radeon naming and remote-shell fixes in guiTOP, a friend's multi-host GPU monitor (#1, #2);
- sampling, compaction and tool-parsing fixes in open-multi-agent.
Predictions are written down before the run and scored after it, and failed predictions are published alongside the ones that held. The receipts are indexed by mechanism in Project Apollo's INDEX.md.
When a published claim turns out to be wrong, the correction goes where the claim was (example). The mistake itself gets an entry in FAILURE_MODES.md.
Earlier RDNA4 bring-up guides (ROCm, PyTorch, ComfyUI, TTS) are in my gists.
I'm looking for a team that pays fairly, treats people well, and cares what its work gets used for. Roles where I'd be useful:
- inference or hardware validation and performance benchmarking;
- solutions or field engineering for AI infrastructure;
- developer relations for inference tools.
Before AI I worked in finance and insurance (F&I) at a car dealership, so explaining trade-offs to people who aren't engineers is familiar ground.
Near Indianapolis, IN. Remote or in person.
Reach me by email at mgalyan@gmail.com or on Discord at apollo_mg.

