Skip to content
transwarp829Public
forked from ggml-org/llama.cpp

About

LLM inference in C/C++

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

11,454 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Important

This is a fork (transwarp829/llama.cpp) that adds an experimental expert pool: an MoE offload mode that keeps a per-layer set of the hottest experts in VRAM and runs the resident (device) and the non-resident (host) expert chains of the same layer as two chains of one graph. Flags, constraints and measured numbers: docs/expert-pool.md. The fork also carries one unmerged upstream ggml change (a -1 expert index skips that MUL_MAT_ID row; the pool uses it as its serve-the-other-chain mask). Everything else is upstream llama.cpp at the commit this fork is based on.

Note

这部分我直接用中文写了。

本fork用于实现专家缓存的概念。目前上游已经出现了思路相近且大概率获得合并的实现,即29887。它所采用的方式,是将未命中的专家通过缓存更新的方式搬运到GPU,然后在GPU上面执行完整的MUL_MAT_ID。

和上游的PR不同,本fork所采用的主要方法是:在GPU侧计算权重缓存命中的专家,在host(CPU)侧原位计算缓存未命中的专家,然后将结果合并到GPU侧,让GPU完成剩余的共享专家/加权归约等等步骤。 为了正确地实现这一功能,上游需要实现对于将同一OP分配到多个后端分别完成计算的功能,但目前来说实现这个功能过于复杂。

因此,本仓库的实现是将MUL_MAT_ID这一单一OP拆分成在各自后端上面分别执行的多个OP,然后合并相加。目前,这依赖于上游的26631 PR。它原本的意义是为可变激活量的模型提供跳过专家计算的选项;在这里被用于将同一算子拆分成两个互补的部分,分别在不同的后端上进行实现。

这涉及到更改计算图的构建,本身会引入额外的复杂性。此外,由于后端调度器根据缓冲区切graph split且不同split之间严格串行的规约,在这种情况下,不同后端的两个MUL_MAT_ID会被串行执行;因此原本的并行语义被破坏,会因此损失速度。

本fork还带有一个更加激进的版本,它通过在后端调度器当中设置一个例外来允许CPU段和GPU段的并行。在我的硬件上进行测试时,这似乎可以带来~5-15%的性能提升。但也可能会因为额外的分拆、驱动提交,和同步成本,而变得更慢。但这个修改同样是对上游行为的重大变更,因此它同样难以被合并。

性能方面上来说,本机的CPU原位计算方案在VRAM预算较为有限的情况下较优,并且不会出现PR在极低VRAM预算下H2D传输需求爆炸、速度坍缩到比纯CPU计算更慢的情况。在VRAM的较高(在我的机器上表现为缓冲区大小>~30%的整层专家权重体积)预算下,上游PR的性能会超越本fork。

从实现或者架构的正确性来说,将同一OP拆在多后端执行的能力,在上游的仓库当中仅属于张量并行所主要使用的meta backend实现,但是它并不允许CPU成为设备的一部分。更高的抽象层,例如RPC服务器倒是允许这种行为,不过目前的效率我猜测只能进行PoC, 无法直接使用。

原位计算类方案的正确演进路径,应该是首先放开CPU参与张量并行,然后再在meta backend内部实现对算力和存储容量有感知的分配机制,并在这一机制内实现池的语义。这会成为本fork的长期探索方向,但不会包含在这一分支当中。

此外,本仓库的缓存池实现不兼容--fit。具体的实现和用法可以在前文提到的md文件当中找到。

llama.cpp

llama

Quick start

A few options to get llama.cpp installed on your machine:

# curl
curl -LsSf https://llama.app/install.sh | sh

# powershell
irm https://llama.app/install.ps1 | iex

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

About

LLM inference in C/C++

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages