Selfhosting and AI nerd from Germany π©πͺ
I play Minecraft, RPGs and survival games, and I build things around them.
- β‘ optllama: a llama.cpp fork that I contribute to heavily, focused on KV cache offload to RAM, unbalanced dual-GPU setups and partial KV cache residency.
- π§ͺ Piggidragon/llama.cpp (branch
llama-tensor): my own optllama fork that merges my open PRs into one usable branch. - π€ Pithagoras: I'm an active contributor to this AI agent project. Image generation and editing with a gallery, voice improvements, UI polish and docs, and most recently the Devices add-on, which lets chats work with files and a shell on paired computers.
- π¦ Pithagoras-Sync: a single Rust binary that lets the Pithagoras agent reach any remote device natively, without SSH.
- π§© OpenWebUI-plugins: tools, functions and pipes for Open WebUI.
- βοΈ Minecraft modding with NeoForge: a magic and dimension mod, plus an Ender Dragon fight overhaul.
| Project | What I contributed |
|---|---|
| generelschwerz/llama.cpp (optllama) | Host-resident KV cache: pipelined delivery, head-split cache under split mode tensor, partial KV residency budget on the slowest link first. Multi-GPU: attention split separate from the tensor split, tied output projection split, meta transport ring per device share. CUDA: quantized-native MMA FlashAttention. Also: DFlash speculative decoding, llama-bench placement and sweep flags, scheduler fixes and the arch-coverage test suite |
| thecodacus/pithagoras | Image generation and editing (several reference pictures, gallery, in-chat viewer), voice fillers, extension screens, animations, the Devices add-on, docs and release fixes |
Ask me about selfhosting, local AI, agents and Minecraft modding.



