Name and Version
$ llama-cli --version
version: 10197 (958d9c0)
built with GNU 15.2.0 for Linux x86_64
Operating systems
Linux
GGML backends
CUDA
Hardware
I have tested this bug on two servers and gotten similar results.
- Server 1: AMD Strix Halo AI MAX 395+ with 128 GB DDR5
- Server 2: AMD 5700G with 64 GB of DDR4
Models
Ternary-Bonsai-27B with the DSpark draft model. In particular, I'm using the Q2_g64 quant of the main model and the Q4_1 quant of the draft model at the URL below:
https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf/tree/main
Problem description & steps to reproduce
I attempt to load the model using llama-server with the following config:
[Ternary-Bonsai-27B]
model = Ternary-Bonsai-27B-Q2_g64.gguf
mmproj = Ternary-Bonsai-27B-mmproj-Q8_0.gguf
model-draft = Ternary-Bonsai-27B-dspark-Q4_1.gguf
ctx-size = 32768
temp = 0.7
top-p = 0.95
top-k = 20
cache-type-v = q8_0
cache-type-k = q8_0
spec-type = draft-dspark
spec-draft-n-max = 2
flash-attn = 1
batch-size = 512
ubatch-size = 512
cache-prompt = 1
parallel = 1
models-max = 1
threads = 6
threads-batch = 6
n-gpu-layers = all
load-mode = mmap
Loading the model fails with the following message:
[39829] 0.00.032.539 I srv load_model: loading model 'Ternary-Bonsai-27B-Q2_g64.gguf'
[39829] 0.00.101.674 E gguf_init_from_reader: tensor 'dspark.fc.weight' has offset 337718592, expected 357584192
[39829] 0.00.101.682 E gguf_init_from_reader: failed to read tensor data
[39829] 0.00.101.777 E llama_model_load: error loading model: llama_model_loader: failed to load model from Ternary-Bonsai-27B-dspark-Q4_1.gguf
[39829] 0.00.101.784 E llama_model_load_from_file_impl: failed to load model
[39829] 0.00.101.798 W srv load_model: [spec] failed to measure draft model memory: failed to load model
First Bad Commit
No response
Relevant log output
Logs
Name and Version
$ llama-cli --version
version: 10197 (958d9c0)
built with GNU 15.2.0 for Linux x86_64
Operating systems
Linux
GGML backends
CUDA
Hardware
I have tested this bug on two servers and gotten similar results.
Models
Ternary-Bonsai-27B with the DSpark draft model. In particular, I'm using the Q2_g64 quant of the main model and the Q4_1 quant of the draft model at the URL below:
https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf/tree/main
Problem description & steps to reproduce
I attempt to load the model using llama-server with the following config:
[Ternary-Bonsai-27B]
model = Ternary-Bonsai-27B-Q2_g64.gguf
mmproj = Ternary-Bonsai-27B-mmproj-Q8_0.gguf
model-draft = Ternary-Bonsai-27B-dspark-Q4_1.gguf
ctx-size = 32768
temp = 0.7
top-p = 0.95
top-k = 20
cache-type-v = q8_0
cache-type-k = q8_0
spec-type = draft-dspark
spec-draft-n-max = 2
flash-attn = 1
batch-size = 512
ubatch-size = 512
cache-prompt = 1
parallel = 1
models-max = 1
threads = 6
threads-batch = 6
n-gpu-layers = all
load-mode = mmap
Loading the model fails with the following message:
[39829] 0.00.032.539 I srv load_model: loading model 'Ternary-Bonsai-27B-Q2_g64.gguf'
[39829] 0.00.101.674 E gguf_init_from_reader: tensor 'dspark.fc.weight' has offset 337718592, expected 357584192
[39829] 0.00.101.682 E gguf_init_from_reader: failed to read tensor data
[39829] 0.00.101.777 E llama_model_load: error loading model: llama_model_loader: failed to load model from Ternary-Bonsai-27B-dspark-Q4_1.gguf
[39829] 0.00.101.784 E llama_model_load_from_file_impl: failed to load model
[39829] 0.00.101.798 W srv load_model: [spec] failed to measure draft model memory: failed to load model
First Bad Commit
No response
Relevant log output
Logs