Description
The latest amperecomputingai/llama.cpp Docker image crashes with Illegal instruction (core dumped) during model warmup on an ARM Neoverse-N1 CPU.
The exact same hardware, model, and llama-server arguments work correctly when using Docker image amperecomputingai/llama.cpp:3.4.2.
This appears to be a regression in a newer Ampere build.
Environment
CPU:
Architecture: aarch64
Vendor ID: ARM
Model name: Neoverse-N1
CPU(s): 4
Features:
fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics
fphp asimdhp cpuid asimdrdm lrcpc dcpop asimddp
/proc/cpuinfo:
CPU implementer : 0x41
CPU architecture: 8
CPU variant : 0x3
CPU part : 0xd0c
CPU revision : 1
Docker sees the same CPU features as the host.
Model:
Qwen3-4B-Q8R16.gguf
Qwen3 4B
Q8R16
3.78 GiB
Reproduction
Using the latest image:
docker run -d --name llama \
-v ~/models:/models \
-p 8080:8080 \
amperecomputingai/llama.cpp:latest \
--model /models/Qwen3-4B-Q8R16.gguf \
--host 0.0.0.0 \
--port 8080 \
--ctx-size 8192
The model loads successfully, the context and KV cache are allocated, but the process crashes as soon as the warmup starts:
system_info: n_threads = 4 (n_threads_batch = 4) / 4 |
CPU : NEON = 1 | ARM_FMA = 1 | FP16_VA = 1 |
MATMUL_INT8 = 1 | DOTPROD = 1 |
LLAMAFILE = 1 | REPACK = 1 | AMPERE = 1 |
...
llama_context: n_ctx = 8192
llama_context: n_batch = 2048
llama_context: n_ubatch = 512
...
common_init_from_params: warming up the model with an empty run - please wait ...
/start.sh: line 3: 14 Illegal instruction (core dumped) /llm/llama-server "$@"
Latest build reports:
build_info: b8896-905ccccf2
Working version
Changing only the image to:
amperecomputingai/llama.cpp:3.4.2
with the same model and command works correctly:
docker run -d --name llama \
-v ~/models:/models \
-p 8080:8080 \
amperecomputingai/llama.cpp:3.4.2 \
--model /models/Qwen3-4B-Q8R16.gguf \
--host 0.0.0.0 \
--port 8080 \
--ctx-size 8192
Relevant startup output:
build: 7834 (943aee2ee) with Clang 19.1.7 for Linux aarch64
system_info: n_threads = 4 (n_threads_batch = 4) / 4 |
CPU : NEON = 1 | ARM_FMA = 1 | FP16_VA = 1 |
DOTPROD = 1 | LLAMAFILE = 1 | REPACK = 1 | AMPERE = 1 |
...
common_init_from_params: warming up the model with an empty run - please wait ...
srv load_model: initializing slots, n_slots = 4
...
main: model loaded
main: server is listening on http://0.0.0.0:8080
main: starting the main loop...
srv update_slots: all slots are idle
srv log_server_r: request: GET /health 127.0.0.1 200
Possibly relevant differences
One visible difference in CPU feature reporting is:
Latest:
3.4.2:
MATMUL_INT8 is not reported
The Neoverse-N1 exposes asimddp / DOTPROD but does not expose i8mm.
I am not assuming MATMUL_INT8 is necessarily the cause, but it may be relevant given that the failure is a SIGILL.
There is also a difference in the model buffer/backend reported during loading.
Latest:
CPU_Mapped model buffer size = 3866.65 MiB
3.4.2:
CPU_AMPERE model buffer size = 3866.65 MiB
Expected behavior
latest should run on Neoverse-N1 / Ampere Altra-class ARMv8.2 hardware as 3.4.2 does, or avoid dispatching instructions unsupported by the detected CPU.
Workaround
Pinning the Docker image to:
amperecomputingai/llama.cpp:3.4.2
avoids the crash.
Please let me know if you need additional verbose logs, a core dump/backtrace, or output from objdump.
Description
The latest
amperecomputingai/llama.cppDocker image crashes withIllegal instruction (core dumped)during model warmup on an ARM Neoverse-N1 CPU.The exact same hardware, model, and llama-server arguments work correctly when using Docker image
amperecomputingai/llama.cpp:3.4.2.This appears to be a regression in a newer Ampere build.
Environment
CPU:
/proc/cpuinfo:Docker sees the same CPU features as the host.
Model:
Reproduction
Using the latest image:
docker run -d --name llama \ -v ~/models:/models \ -p 8080:8080 \ amperecomputingai/llama.cpp:latest \ --model /models/Qwen3-4B-Q8R16.gguf \ --host 0.0.0.0 \ --port 8080 \ --ctx-size 8192The model loads successfully, the context and KV cache are allocated, but the process crashes as soon as the warmup starts:
Latest build reports:
Working version
Changing only the image to:
with the same model and command works correctly:
docker run -d --name llama \ -v ~/models:/models \ -p 8080:8080 \ amperecomputingai/llama.cpp:3.4.2 \ --model /models/Qwen3-4B-Q8R16.gguf \ --host 0.0.0.0 \ --port 8080 \ --ctx-size 8192Relevant startup output:
Possibly relevant differences
One visible difference in CPU feature reporting is:
Latest:
3.4.2:
The Neoverse-N1 exposes
asimddp/ DOTPROD but does not exposei8mm.I am not assuming
MATMUL_INT8is necessarily the cause, but it may be relevant given that the failure is aSIGILL.There is also a difference in the model buffer/backend reported during loading.
Latest:
3.4.2:
Expected behavior
latestshould run on Neoverse-N1 / Ampere Altra-class ARMv8.2 hardware as3.4.2does, or avoid dispatching instructions unsupported by the detected CPU.Workaround
Pinning the Docker image to:
avoids the crash.
Please let me know if you need additional verbose logs, a core dump/backtrace, or output from
objdump.