Important
This is a fork (transwarp829/llama.cpp) that adds an experimental expert
pool: an MoE offload mode that keeps a per-layer set of the hottest experts in
VRAM and runs the resident (device) and the non-resident (host) expert chains of
the same layer as two chains of one graph. Flags, constraints and measured
numbers: docs/expert-pool.md. The fork also carries one
unmerged upstream ggml change (a -1 expert index skips that MUL_MAT_ID row;
the pool uses it as its serve-the-other-chain mask). Everything else is upstream
llama.cpp at the commit this fork is based on.
Note
这部分我直接用中文写了。
本fork用于实现专家缓存的概念。目前上游已经出现了思路相近且大概率获得合并的实现,即29887。它所采用的方式,是将未命中的专家通过缓存更新的方式搬运到GPU,然后在GPU上面执行完整的MUL_MAT_ID。
和上游的PR不同,本fork所采用的主要方法是:在GPU侧计算权重缓存命中的专家,在host(CPU)侧原位计算缓存未命中的专家,然后将结果合并到GPU侧,让GPU完成剩余的共享专家/加权归约等等步骤。 为了正确地实现这一功能,上游需要实现对于将同一OP分配到多个后端分别完成计算的功能,但目前来说实现这个功能过于复杂。
因此,本仓库的实现是将MUL_MAT_ID这一单一OP拆分成在各自后端上面分别执行的多个OP,然后合并相加。目前,这依赖于上游的26631 PR。它原本的意义是为可变激活量的模型提供跳过专家计算的选项;在这里被用于将同一算子拆分成两个互补的部分,分别在不同的后端上进行实现。
这涉及到更改计算图的构建,本身会引入额外的复杂性。此外,由于后端调度器根据缓冲区切graph split且不同split之间严格串行的规约,在这种情况下,不同后端的两个MUL_MAT_ID会被串行执行;因此原本的并行语义被破坏,会因此损失速度。
本fork还带有一个更加激进的版本,它通过在后端调度器当中设置一个例外来允许CPU段和GPU段的并行。在我的硬件上进行测试时,这似乎可以带来~5-15%的性能提升。但也可能会因为额外的分拆、驱动提交,和同步成本,而变得更慢。但这个修改同样是对上游行为的重大变更,因此它同样难以被合并。
性能方面上来说,本机的CPU原位计算方案在VRAM预算较为有限的情况下较优,并且不会出现PR在极低VRAM预算下H2D传输需求爆炸、速度坍缩到比纯CPU计算更慢的情况。在VRAM的较高(在我的机器上表现为缓冲区大小>~30%的整层专家权重体积)预算下,上游PR的性能会超越本fork。
从实现或者架构的正确性来说,将同一OP拆在多后端执行的能力,在上游的仓库当中仅属于张量并行所主要使用的meta backend实现,但是它并不允许CPU成为设备的一部分。更高的抽象层,例如RPC服务器倒是允许这种行为,不过目前的效率我猜测只能进行PoC, 无法直接使用。
原位计算类方案的正确演进路径,应该是首先放开CPU参与张量并行,然后再在meta backend内部实现对算力和存储容量有感知的分配机制,并在这一机制内实现池的语义。这会成为本fork的长期探索方向,但不会包含在这一分支当中。
此外,本仓库的缓存池实现不兼容--fit。具体的实现和用法可以在前文提到的md文件当中找到。
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
# curl
curl -LsSf https://llama.app/install.sh | sh
# powershell
irm https://llama.app/install.ps1 | iex- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with llama cli
|
Built-in web UI against llama serve
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

