Summary
Attempting to load any model from the prism-ml/Bonsai-*-gguf family
(1-bit LLMs by PrismML) fails with the following error:
Failed to load model: Failed to load model even at minimum context (2048).
This may indicate insufficient memory, a corrupted model file, or an
unsupported model format. (Failed to load model)
Device
- Samsung Galaxy S23 (Snapdragon 8 Gen 2, 8GB RAM)
- Android 14
- Off Grid (latest Play Store version)
Steps to Reproduce
- Models → Search "Bonsai"
- Tap
Bonsai-4B-gguf or Bonsai-8B-gguf by prism-ml
- Download and attempt to load
- Error appears immediately on load
Root Cause
The Bonsai models use a Q1_0_g128 quantization format — a native 1-bit format (1 bit per weight + 1 FP16 scale per 128-weight group) introduced by PrismML in their own llama.cpp fork: https://github.com/PrismML-Eng/llama.cpp
This format is not yet merged into upstream llama.cpp, which is what llama.rn currently bundles.
Why This Matters
Bonsai 4B is only ~1 GB and benchmarks comparably to standard 8B models. It would be an ideal fit for mobile — but requires the Q1_0_g128 kernels to load.
Request
Could you update llama.rn / the bundled llama.cpp to include PrismML's Q1_0_g128 support? Their fork is MIT-licensed and the relevant changes appear to be isolated to the new quant type kernels.
Reference:
Summary
Attempting to load any model from the
prism-ml/Bonsai-*-gguffamily(1-bit LLMs by PrismML) fails with the following error:
Device
Steps to Reproduce
Bonsai-4B-gguforBonsai-8B-ggufbyprism-mlRoot Cause
The Bonsai models use a Q1_0_g128 quantization format — a native 1-bit format (1 bit per weight + 1 FP16 scale per 128-weight group) introduced by PrismML in their own llama.cpp fork: https://github.com/PrismML-Eng/llama.cpp
This format is not yet merged into upstream llama.cpp, which is what
llama.rncurrently bundles.Why This Matters
Bonsai 4B is only ~1 GB and benchmarks comparably to standard 8B models. It would be an ideal fit for mobile — but requires the Q1_0_g128 kernels to load.
Request
Could you update
llama.rn/ the bundledllama.cppto include PrismML's Q1_0_g128 support? Their fork is MIT-licensed and the relevant changes appear to be isolated to the new quant type kernels.Reference: