DwarfStar is a native engine for a few large MoE models (DeepSeek V4/V4.1
Flash, GLM 5.x Flash, Qwen3.8-Flash-Next) in its own GGUF layouts, with
first-class Strix Halo ROCm support and SSD expert streaming. It runs as a
child of 1bit serve, like ZINC.
- third_party/ds4 pins antirez/ds4 0aaea5a238fb (MIT); bump-ds4.yml follows
upstream main daily.
- scripts/build-ds4.sh builds a copy of the tree (the submodule stays clean)
for rocm, cuda, metal or cpu. For ROCm it finds TheRock's SDK and passes its
include/ with -isystem: clang otherwise searches it after /usr/include, where
a distro HIP of another version shadows it. DS4_TEST=1 runs DwarfStar's
model-free routed-MoE GPU test.
- ONEBIT_DS4 builds it; serve --device ds4 starts ds4-server (--ctx-size,
--ssd-streaming, --ds4 PATH). A child now carries its own readiness path:
ds4-server has no /health and opens its port once the model is loaded, so
serve waits on its /v1/models.
- docs/dwarfstar.md; serve.md, README and NOTICE updated.
Verified on Strix Halo (TheRock ROCm): builds, routed-MoE test PASS, engine
ctest 7/7, the missing-model and --ssd-streaming error paths. A full model
(DeepSeek V4 Flash Q2, 81 GiB) is still to run: no room on the box.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DwarfStar (MIT) is a native engine for a few large MoE models (DeepSeek V4/V4.1 Flash, GLM 5.x Flash, Qwen3.8-Flash-Next) in its own GGUF layouts, with first-class Strix Halo ROCm support and SSD expert streaming. Its GGUFs are the most downloaded DeepSeek V4.1 Flash files on HF (
antirez/deepseek-v4.1-flash-gguf, 940k), and llama.cpp cannot load them. This adds it as a child backend, like ZINC.third_party/ds4: antirez/ds40aaea5a238fb;bump-ds4.ymlfollows upstream main daily.scripts/build-ds4.sh: builds a copy of the tree (the submodule stays clean) for rocm / cuda / metal / cpu. ROCm finds TheRock's SDK and passes itsinclude/with-isystem(clang otherwise searches it after/usr/include, where a distro HIP of another version shadows it).DS4_TEST=1runs DwarfStar's model-free routed-MoE GPU test.-DONEBIT_DS4=ON;1bit serve --device ds4with--ctx-size,--ssd-streaming,--ds4 PATH. A child now carries its own readiness path:ds4-serverhas no/healthand opens its port once the model is loaded, soservewaits on/v1/models.docs/dwarfstar.md; serve.md, README (repo XDNA: pin upstream amd/xdna-driver (and its XRT), build it privately, keep it current #12) and NOTICE.Verified on Strix Halo (TheRock ROCm):
DS4_TEST=1 scripts/build-ds4.sh build/ds4 rocm-DONEBIT_DS4=ONserve --device ds4 -m <missing>--ssd-streamingon another deviceNot yet run: a full model. DeepSeek V4 Flash Q2 is 81 GiB on disk and needs about that much free memory; the box has 89 GB of disk and ~60 GB of memory free while other work runs. Hold the merge until it answers a chat through
serve --device ds4.🤖 Generated with Claude Code