0.21.532.917 I srv load: spawning server instance with name=Qwen3.6-35B-A3B-MTP-GGUF_Q8_K_XL on port 49457
0.21.532.942 I srv load: spawning server instance with args:
0.21.532.943 I srv load: /usr/bin/llama-server
0.21.532.943 I srv load: --cache-prompt
0.21.532.943 I srv load: --cache-reuse
0.21.532.944 I srv load: 256
0.21.532.944 I srv load: --host
0.21.532.944 I srv load: 127.0.0.1
0.21.532.944 I srv load: --jinja
0.21.532.944 I srv load: --min-p
0.21.532.944 I srv load: 0.01
0.21.532.946 I srv load: --no-mmap
0.21.532.946 I srv load: --numa
0.21.532.946 I srv load: numactl
0.21.532.946 I srv load: --poll
0.21.532.946 I srv load: 50
0.21.532.946 I srv load: --port
0.21.532.946 I srv load: 49457
0.21.532.947 I srv load: --prio
0.21.532.947 I srv load: 2
0.21.532.947 I srv load: --sleep-idle-seconds
0.21.532.947 I srv load: 3600
0.21.532.947 I srv load: --spec-draft-n-max
0.21.532.948 I srv load: 3
0.21.532.948 I srv load: --spec-draft-n-min
0.21.532.948 I srv load: 1
0.21.532.948 I srv load: --draft-p-min
0.21.532.948 I srv load: 0.9
0.21.532.948 I srv load: --spec-ngram-map-k-size-m
0.21.532.948 I srv load: 4
0.21.532.952 I srv load: --spec-ngram-map-k-size-n
0.21.532.953 I srv load: 8
0.21.532.953 I srv load: --spec-type
0.21.532.953 I srv load: ngram-map-k
0.21.532.953 I srv load: --swa-full
0.21.532.953 I srv load: --temperature
0.21.532.954 I srv load: 0.7
0.21.532.954 I srv load: --top-k
0.21.532.954 I srv load: 20
0.21.532.954 I srv load: --top-p
0.21.532.954 I srv load: 0.95
0.21.532.954 I srv load: --no-warmup
0.21.532.954 I srv load: --alias
0.21.532.955 I srv load: Qwen3.6-35B-A3B-MTP-GGUF_Q8_K_XL
0.21.532.955 I srv load: --batch-size
0.21.532.955 I srv load: 1024
0.21.532.956 I srv load: --ctx-size
0.21.532.956 I srv load: 256000
0.21.532.956 I srv load: --cache-ram
0.21.532.956 I srv load: 32768
0.21.532.956 I srv load: --cache-type-k
0.21.532.956 I srv load: q8_0
0.21.532.957 I srv load: --cache-type-v
0.21.532.957 I srv load: f16
0.21.532.958 I srv load: --swa-checkpoints
0.21.532.958 I srv load: 32
0.21.532.958 I srv load: --direct-io
0.21.532.958 I srv load: --flash-attn
0.21.532.958 I srv load: on
0.21.532.959 I srv load: --fit
0.21.532.959 I srv load: off
0.21.532.959 I srv load: --hf-repo
0.21.532.959 I srv load: unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q8_K_XL
0.21.532.959 I srv load: --kv-offload
0.21.532.959 I srv load: --no-kv-unified
0.21.532.960 I srv load: --n-gpu-layers
0.21.532.960 I srv load: 99
0.21.532.960 I srv load: --parallel
0.21.532.960 I srv load: 2
0.21.532.960 I srv load: --split-mode
0.21.532.961 I srv load: none
0.21.532.961 I srv load: --slot-prompt-similarity
0.21.532.961 I srv load: 0.75
0.21.532.961 I srv load: --threads
0.21.532.962 I srv load: 8
0.21.532.962 I srv load: --threads-batch
0.21.532.962 I srv load: 12
0.21.532.962 I srv load: --ubatch-size
0.21.532.962 I srv load: 1024
[49457] 3.26.682.896 I slot print_timing: id 1 | task 104 | prompt processing, n_tokens = 12276, progress = 0.90, t = 29.44 s / 417.05 tokens per second
[49457] 3.28.387.929 I slot print_timing: id 1 | task 104 | prompt processing, n_tokens = 12654, progress = 0.93, t = 31.14 s / 406.35 tokens per second
[49457] 3.32.256.559 I slot print_timing: id 0 | task 1 | n_decoded = 196, tg = 1.27 t/s
[49457] 3.32.256.611 I slot print_timing: id 1 | task 104 | prompt processing, n_tokens = 13628, progress = 1.00, t = 35.01 s / 389.27 tokens per second
[49457] 3.32.265.192 I slot create_check: id 1 | task 104 | created context checkpoint 1 of 32 (pos_min = 13627, pos_max = 13627, n_tokens = 13628, size = 62.813 MiB)
[49457] 3.32.596.646 I slot print_timing: id 1 | task 104 | prompt processing, n_tokens = 13674, progress = 1.00, t = 35.35 s / 386.83 tokens per second
[49457] 3.32.720.244 I common_ngram_map_begin: refresh map: idx_last_draft=23242, new begin=13678, #keys_checked=10, #keys_del=0, #values_del=0, #hashes_upd=5
[49457] 3.33.598.926 I reasoning-budget: deactivated (natural end)
[49457] /usr/include/c++/16.1.1/bits/stl_vector.h:1272: constexpr std::vector<_Tp, _Alloc>::const_reference std::vector<_Tp, _Alloc>::operator[](size_type) const [with _Tp = int; _Alloc = std::allocator<int>; const_reference = const int&; size_type = long unsigned int]: Assertion '__n < this->size()' failed.
(gdb) bt full
#0 __pthread_kill_implementation (threadid=<optimized out>, signo=signo@entry=6, no_tid=no_tid@entry=0) at pthread_kill.c:44
tid = <optimized out>
ret = 0
pd = <optimized out>
old_mask = {__val = {94486818072208}}
ret = <optimized out>
#1 0x00007f3ffde9a363 in __pthread_kill_internal (threadid=<optimized out>, signo=6) at pthread_kill.c:89
No locals.
#2 0x00007f3ffde3e7d0 in __GI_raise (sig=sig@entry=6) at ../sysdeps/posix/raise.c:26
ret = <optimized out>
#3 0x00007f3ffde25681 in __GI_abort () at abort.c:77
act = {__sigaction_handler = {sa_handler = 0x4f8, sa_sigaction = 0x4f8}, sa_mask = {__val = {139912816337128, 3, 24, 94486902120320, 94486902120344, 94486902120344, 24, 140720762873856, 94486902118496, 94486902120376, 94486902120376, 94486899295216,
94486899295216, 94486899295264, 140720762873824, 48}}, sa_flags = 0, sa_restorer = 0x7ffc1b156730}
#4 0x00007f3ffca9d5b9 in std::__glibcxx_assert_fail (file=<optimized out>, line=<optimized out>, function=<optimized out>, condition=<optimized out>) at /usr/src/debug/gcc/gcc/libstdc++-v3/src/c++11/assert_fail.cc:41
No locals.
#5 0x00007f3ffd8aafbb in ?? () from /usr/lib/libllama-common.so.0
No symbol table info available.
#6 0x00007f3ffda9b5d6 in ?? () from /usr/lib/libllama-common.so.0
No symbol table info available.
#7 0x00007f3ffdaa6e77 in common_speculative_draft(common_speculative*) () from /usr/lib/libllama-common.so.0
No symbol table info available.
#8 0x00007f3ffe33c446 in ?? () from /usr/lib/libllama-server-impl.so
No symbol table info available.
#9 0x00007f3ffe3dca33 in server_queue::start_loop(long) () from /usr/lib/libllama-server-impl.so
No symbol table info available.
#10 0x00007f3ffe2c346c in llama_server(int, char**) () from /usr/lib/libllama-server-impl.so
No symbol table info available.
#11 0x00007f3ffde27741 in __libc_start_call_main (main=main@entry=0x55ef3406d020, argc=argc@entry=77, argv=argv@entry=0x7ffc1b15f478) at ../sysdeps/nptl/libc_start_call_main.h:59
self = <optimized out>
result = <optimized out>
unwind_buf = {cancel_jmp_buf = {{jmp_buf = {0, 6630328609852258281, 140720762909816, 77, 139912832815104, 94485858418096, 6630328609864841193, 6738622348876871657}, mask_was_saved = 0}}, priv = {pad = {0x0, 0x0, 0x0, 0x7ffc1b15f478}, data = {prev = 0x0,
cleanup = 0x0, canceltype = 0}}}
not_first_call = <optimized out>
#12 0x00007f3ffde27879 in __libc_start_main_impl (main=0x55ef3406d020, argc=77, argv=0x7ffc1b15f478, init=<optimized out>, fini=<optimized out>, rtld_fini=<optimized out>, stack_end=0x7ffc1b15f468) at ../csu/libc-start.c:360
No locals.
#13 0x000055ef3406d055 in ?? ()
No symbol table info available.
(gdb)
Name and Version
$ llama-server --version
version: 9413 (24d6d7c)
built with GNU 16.1.1 for Linux x86_64
Operating systems
Linux
GGML backends
HIP
Hardware
AMD RYZEN AI MAX+ PRO 395 w/ Radeon 8060S
Models
Qwen3.6-35B-A3B-MTP-GGUF_Q8_K_XL
Problem description & steps to reproduce
From time to time llama-server crashes during prompt-eval. Usually takes 20-30 minutes to reproduce the issue.
Server is launched via "/usr/bin/numactl --cpunodebind=0 --membind=0 /usr/bin/llama-server --numa numactl ......" using a models-preset file.
--models-max is set to 1.
Going to make a new build with debug-symbols and run a test start of next week and hopefully have a more informative stacktrace.
First Bad Commit
Somewhere between b9413 and b9371 this started appearing. Have not yet had the time to perform a bisect.
Relevant log output
Logs