Skip to content
View apollo-mg's full-sized avatar
  • Self Employed
  • Near Indianapolis Indiana

Block or report apollo-mg

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
apollo-mg/README.md

Mark Galyan

Systems engineer working on local LLM inference: measuring how models actually behave on specific hardware, and fixing it upstream when they don't behave. My lab is an RX 9070 XT (RDNA4) desktop plus six Tesla P100s in two home servers.

I use AI to make AI, break AI, and make it again. Every day. The agents (Claude, Gemini) write the code; the hardware, the experiments, the review and the final calls are mine. Every claim below links to its evidence.

Highlights

  • KV caches that survive a restart. llama-server restored a saved session and then discarded it on the next request. A 117-line checkpoint sidecar fixed it: a parked 100K-token session comes back in about 4 s instead of a 12-minute re-prefill. Merged in llama-cpp-turboquant #206. Writeup
  • Tesla P100s were doing quality-sensitive math in fp16 for years. A three-line fix moved median KLD against an fp32 reference from 0.0023 to 0.000001, with no speed cost. Merged in llama-cpp-turboquant #212 and buun-llama-cpp #80, and reported upstream as ggml-org/llama.cpp #25593. Writeup
  • Tool calling restored for a Qwen3.8 fine-tune. The chat template it shipped with silently dropped tool_calls from the conversation history. The model's author adopted the fixed template (discussion). Template and test report
  • CUDA-on-AMD defects on RDNA4. Reported two defects in SCALE on gfx1201:
    • cudaMemGetInfo charges 4.00x the requested bytes (#67);
    • cuModuleGetFunction returns success with an unusable handle (#68).
  • Other merged work:
    • AMD Radeon naming and remote-shell fixes in guiTOP, a friend's multi-host GPU monitor (#1, #2);
    • sampling, compaction and tool-parsing fixes in open-multi-agent.

How I work

Predictions are written down before the run and scored after it, and failed predictions are published alongside the ones that held. The receipts are indexed by mechanism in Project Apollo's INDEX.md.

When a published claim turns out to be wrong, the correction goes where the claim was (example). The mistake itself gets an entry in FAILURE_MODES.md.

Earlier RDNA4 bring-up guides (ROCm, PyTorch, ComfyUI, TTS) are in my gists.

Open to work

I'm looking for a team that pays fairly, treats people well, and cares what its work gets used for. Roles where I'd be useful:

  • inference or hardware validation and performance benchmarking;
  • solutions or field engineering for AI infrastructure;
  • developer relations for inference tools.

Before AI I worked in finance and insurance (F&I) at a car dealership, so explaining trade-offs to people who aren't engineers is familiar ground.

Near Indianapolis, IN. Remote or in person. Reach me by email at mgalyan@gmail.com or on Discord at apollo_mg.

Popular repositories Loading

  1. Project-Apollo Project-Apollo Public

    Home lab for local LLM inference on RDNA4 and Tesla P100: preregistered experiments, receipts, upstream fixes.

    Python 3 1

  2. ComfyUI-Zluda-MyFix ComfyUI-Zluda-MyFix Public

    My working ComfyUI ZLUDA

    Python 1 1

  3. RDNA4-ComfyUI-Patches RDNA4-ComfyUI-Patches Public

    Python 1

  4. urban-spork urban-spork Public

  5. apollo-mg apollo-mg Public

    Profile README

  6. open-multi-agent open-multi-agent Public

    Forked from open-multi-agent/open-multi-agent

    TypeScript multi-agent orchestration engine — one runTeam() call from goal to result. Multi-model teams, auto task decomposition, parallel execution. 3 runtime dependencies.

    TypeScript