Skip to content
View Piggidragon's full-sized avatar
  • Germany

Highlights

  • Pro

Block or report Piggidragon

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Piggidragon/README.md

Hi, I'm Piggidragon πŸ‰

Selfhosting and AI nerd from Germany πŸ‡©πŸ‡ͺ
I play Minecraft, RPGs and survival games, and I build things around them.

Followers Location: Germany


πŸ”­ What I'm working on

  • ⚑ optllama: a llama.cpp fork that I contribute to heavily, focused on KV cache offload to RAM, unbalanced dual-GPU setups and partial KV cache residency.
  • πŸ§ͺ Piggidragon/llama.cpp (branch llama-tensor): my own optllama fork that merges my open PRs into one usable branch.
  • πŸ€– Pithagoras: I'm an active contributor to this AI agent project. Image generation and editing with a gallery, voice improvements, UI polish and docs, and most recently the Devices add-on, which lets chats work with files and a shell on paired computers.
  • πŸ¦€ Pithagoras-Sync: a single Rust binary that lets the Pithagoras agent reach any remote device natively, without SSH.
  • 🧩 OpenWebUI-plugins: tools, functions and pipes for Open WebUI.
  • ⛏️ Minecraft modding with NeoForge: a magic and dimension mod, plus an Ender Dragon fight overhaul.

πŸ› οΈ Tech stack

Python Rust Java JavaScript LaTeX

🀝 Open source contributions

Project What I contributed
generelschwerz/llama.cpp (optllama) Host-resident KV cache: pipelined delivery, head-split cache under split mode tensor, partial KV residency budget on the slowest link first. Multi-GPU: attention split separate from the tensor split, tied output projection split, meta transport ring per device share. CUDA: quantized-native MMA FlashAttention. Also: DFlash speculative decoding, llama-bench placement and sweep flags, scheduler fixes and the arch-coverage test suite
thecodacus/pithagoras Image generation and editing (several reference pictures, gallery, in-chat viewer), voice fillers, extension screens, animations, the Devices add-on, docs and release fixes

πŸ“Š GitHub stats

GitHub stats Top languages

πŸ’¬ Find me on Discord

Discord: Piggidragon optllama Discord Codacus Discord

Ask me about selfhosting, local AI, agents and Minecraft modding.

Pinned Loading

  1. Pithagoras-Sync Pithagoras-Sync Public

    A single rust binary for the Pithagoras Agent to access any remote device natively without SSH.

    Rust 1

  2. thecodacus/pithagoras thecodacus/pithagoras Public

    Pithagoras β€” hosted web UI for the pi coding agent. Give it a task, close the browser, come back when it's done.

    TypeScript 275 75

  3. GenerelSchwerz/llama.cpp GenerelSchwerz/llama.cpp Public

    Forked from ggml-org/llama.cpp

    LLM inference in C/C++

    C++ 138 24

  4. llama.cpp llama.cpp Public

    Forked from GenerelSchwerz/llama.cpp

    LLM inference in C/C++

    C++ 1