Skip to content

Illegal instruction on Neoverse-N1 with latest Docker image; 3.4.2 works #21

Description

@BitRevenant

Description

The latest amperecomputingai/llama.cpp Docker image crashes with Illegal instruction (core dumped) during model warmup on an ARM Neoverse-N1 CPU.

The exact same hardware, model, and llama-server arguments work correctly when using Docker image amperecomputingai/llama.cpp:3.4.2.

This appears to be a regression in a newer Ampere build.

Environment

CPU:

Architecture:        aarch64
Vendor ID:           ARM
Model name:          Neoverse-N1
CPU(s):              4

Features:
fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics
fphp asimdhp cpuid asimdrdm lrcpc dcpop asimddp

/proc/cpuinfo:

CPU implementer : 0x41
CPU architecture: 8
CPU variant     : 0x3
CPU part        : 0xd0c
CPU revision    : 1

Docker sees the same CPU features as the host.

Model:

Qwen3-4B-Q8R16.gguf
Qwen3 4B
Q8R16
3.78 GiB

Reproduction

Using the latest image:

docker run -d --name llama \
  -v ~/models:/models \
  -p 8080:8080 \
  amperecomputingai/llama.cpp:latest \
  --model /models/Qwen3-4B-Q8R16.gguf \
  --host 0.0.0.0 \
  --port 8080 \
  --ctx-size 8192

The model loads successfully, the context and KV cache are allocated, but the process crashes as soon as the warmup starts:

system_info: n_threads = 4 (n_threads_batch = 4) / 4 |
CPU : NEON = 1 | ARM_FMA = 1 | FP16_VA = 1 |
MATMUL_INT8 = 1 | DOTPROD = 1 |
LLAMAFILE = 1 | REPACK = 1 | AMPERE = 1 |

...

llama_context: n_ctx         = 8192
llama_context: n_batch       = 2048
llama_context: n_ubatch      = 512

...

common_init_from_params: warming up the model with an empty run - please wait ...
/start.sh: line 3: 14 Illegal instruction (core dumped) /llm/llama-server "$@"

Latest build reports:

build_info: b8896-905ccccf2

Working version

Changing only the image to:

amperecomputingai/llama.cpp:3.4.2

with the same model and command works correctly:

docker run -d --name llama \
  -v ~/models:/models \
  -p 8080:8080 \
  amperecomputingai/llama.cpp:3.4.2 \
  --model /models/Qwen3-4B-Q8R16.gguf \
  --host 0.0.0.0 \
  --port 8080 \
  --ctx-size 8192

Relevant startup output:

build: 7834 (943aee2ee) with Clang 19.1.7 for Linux aarch64

system_info: n_threads = 4 (n_threads_batch = 4) / 4 |
CPU : NEON = 1 | ARM_FMA = 1 | FP16_VA = 1 |
DOTPROD = 1 | LLAMAFILE = 1 | REPACK = 1 | AMPERE = 1 |

...

common_init_from_params: warming up the model with an empty run - please wait ...

srv load_model: initializing slots, n_slots = 4

...

main: model loaded
main: server is listening on http://0.0.0.0:8080
main: starting the main loop...
srv update_slots: all slots are idle
srv log_server_r: request: GET /health 127.0.0.1 200

Possibly relevant differences

One visible difference in CPU feature reporting is:

Latest:

MATMUL_INT8 = 1

3.4.2:

MATMUL_INT8 is not reported

The Neoverse-N1 exposes asimddp / DOTPROD but does not expose i8mm.

I am not assuming MATMUL_INT8 is necessarily the cause, but it may be relevant given that the failure is a SIGILL.

There is also a difference in the model buffer/backend reported during loading.

Latest:

CPU_Mapped model buffer size = 3866.65 MiB

3.4.2:

CPU_AMPERE model buffer size = 3866.65 MiB

Expected behavior

latest should run on Neoverse-N1 / Ampere Altra-class ARMv8.2 hardware as 3.4.2 does, or avoid dispatching instructions unsupported by the detected CPU.

Workaround

Pinning the Docker image to:

amperecomputingai/llama.cpp:3.4.2

avoids the crash.

Please let me know if you need additional verbose logs, a core dump/backtrace, or output from objdump.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions