368.34.643.407 W slot update_slots: id 0 | task 145480 | 29 198 248059 248046
368.34.643.407 W slot update_slots: id 0 | task 145480 | 29 198 248059 248046
368.34.643.409 W slot update_slots: id 0 | task 145480 | n_past = 12314, slot.prompt.tokens.size() = 12314, seq_id = 0, pos_min = 12313, n_swa = 0
368.34.643.410 I slot update_slots: id 0 | task 145480 | Checking checkpoint with [12218, 12218] against 12313...
368.34.657.656 W slot update_slots: id 0 | task 145480 | restored context checkpoint (pos_min = 12218, pos_max = 12218, n_tokens = 12219, n_past = 12219, size = 171.487 MiB)
368.34.894.607 I slot create_check: id 0 | task 145480 | created context checkpoint 5 of 32 (pos_min = 12741, pos_max = 12741, n_tokens = 12742, size = 172.423 MiB)
368.35.241.319 I reasoning-budget: deactivated (natural end)
368.35.866.875 I slot print_timing: id 0 | task 145480 | n_decoded = 103, tg = 110.13 t/s
368.37.090.834 I slot print_timing: id 0 | task 145480 | prompt eval time = 288.08 ms / 527 tokens ( 0.55 ms per token, 1829.35 tokens per second)
368.37.090.838 I slot print_timing: id 0 | task 145480 | eval time = 2159.22 ms / 254 tokens ( 8.50 ms per token, 117.64 tokens per second)
368.37.090.839 I slot print_timing: id 0 | task 145480 | total time = 2447.30 ms / 781 tokens
368.37.090.840 I slot print_timing: id 0 | task 145480 | graphs reused = 52546
368.37.090.841 I slot print_timing: id 0 | task 145480 | draft acceptance = 0.99505 ( 201 accepted / 202 generated)
368.37.090.855 I statistics draft-mtp: #calls(b,g,a) = 566 142602 107915, #gen drafts = 107915, #acc drafts = 103668, #gen tokens = 294686, #acc tokens = 274174, dur(b,g,a) = 0.412, 851457.082, 118.639 ms
368.37.091.409 I slot release: id 0 | task 145480 | stop processing: n_tokens = 13002, truncated = 0
368.37.091.461 I srv update_slots: all slots are idle
368.37.263.055 I srv params_from_: Chat format: peg-native
368.37.264.238 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.371 (> 0.100 thold), f_keep = 1.000
368.37.265.044 I reasoning-budget: activated, budget=2147483647 tokens
368.37.265.250 I slot launch_slot_: id 0 | task 145539 | processing task, is_child = 0
368.37.265.276 W slot update_slots: id 0 | task 145539 | old: ...
</tool_call><|im_end|>
| <|endoftext|>
368.37.265.276 W slot update_slots: id 0 | task 145539 | new: ...
</tool_call><|im_end|>
| <|im_start|>
368.37.265.276 W slot update_slots: id 0 | task 145539 | 198 248059 248046 198 248044
368.37.265.277 W slot update_slots: id 0 | task 145539 | 198 248059 248046 198 248045
368.37.265.279 W slot update_slots: id 0 | task 145539 | n_past = 13001, slot.prompt.tokens.size() = 13002, seq_id = 0, pos_min = 13001, n_swa = 0
368.37.265.279 I slot update_slots: id 0 | task 145539 | Checking checkpoint with [12741, 12741] against 13000...
368.37.278.750 W slot update_slots: id 0 | task 145539 | restored context checkpoint (pos_min = 12741, pos_max = 12741, n_tokens = 12742, n_past = 12742, size = 172.423 MiB)
368.40.199.291 I slot update_slots: id 0 | task 145539 | 8192 tokens since last checkpoint at 12742, creating new checkpoint during processing at position 22982
368.40.226.084 I slot create_check: id 0 | task 145539 | created context checkpoint 6 of 32 (pos_min = 20933, pos_max = 20933, n_tokens = 20934, size = 187.079 MiB)
368.40.991.434 I slot print_timing: id 0 | task 145539 | prompt processing, n_tokens = 10240, progress = 0.66, t = 3.73 s / 2748.13 tokens per second
368.41.770.369 I slot print_timing: id 0 | task 145539 | prompt processing, n_tokens = 12288, progress = 0.71, t = 4.51 s / 2727.57 tokens per second
368.42.564.058 I slot print_timing: id 0 | task 145539 | prompt processing, n_tokens = 14336, progress = 0.77, t = 5.30 s / 2705.52 tokens per second
368.43.372.331 I slot print_timing: id 0 | task 145539 | prompt processing, n_tokens = 16384, progress = 0.83, t = 6.11 s / 2682.79 tokens per second
368.43.372.546 I slot update_slots: id 0 | task 145539 | 8192 tokens since last checkpoint at 20934, creating new checkpoint during processing at position 31174
368.43.398.955 I slot create_check: id 0 | task 145539 | created context checkpoint 7 of 32 (pos_min = 29125, pos_max = 29125, n_tokens = 29126, size = 201.735 MiB)
368.44.218.959 I slot print_timing: id 0 | task 145539 | prompt processing, n_tokens = 18432, progress = 0.89, t = 6.95 s / 2650.68 tokens per second
368.45.050.753 I slot print_timing: id 0 | task 145539 | prompt processing, n_tokens = 20480, progress = 0.95, t = 7.79 s / 2630.54 tokens per second
368.45.590.574 I slot print_timing: id 0 | task 145539 | prompt processing, n_tokens = 21753, progress = 0.99, t = 8.33 s / 2612.88 tokens per second
368.45.619.382 I slot create_check: id 0 | task 145539 | created context checkpoint 8 of 32 (pos_min = 34494, pos_max = 34494, n_tokens = 34495, size = 211.341 MiB)
368.45.835.055 I slot print_timing: id 0 | task 145539 | prompt processing, n_tokens = 22265, progress = 1.00, t = 8.57 s / 2598.08 tokens per second
368.45.865.881 I slot create_check: id 0 | task 145539 | created context checkpoint 9 of 32 (pos_min = 35006, pos_max = 35006, n_tokens = 35007, size = 212.257 MiB)
368.49.934.350 I slot print_timing: id 0 | task 145539 | n_decoded = 100, tg = 24.81 t/s
368.52.965.132 I slot print_timing: id 0 | task 145539 | n_decoded = 175, tg = 24.78 t/s
368.55.974.722 I slot print_timing: id 0 | task 145539 | n_decoded = 249, tg = 24.73 t/s
...
392.06.145.199 I slot print_timing: id 0 | task 145539 | n_decoded = 31920, tg = 22.80 t/s
392.09.171.862 I slot print_timing: id 0 | task 145539 | n_decoded = 31984, tg = 22.79 t/s
392.09.927.219 I slot print_timing: id 0 | task 145539 | prompt eval time = 8638.69 ms / 22269 tokens ( 0.39 ms per token, 2577.82 tokens per second)
392.09.927.224 I slot print_timing: id 0 | task 145539 | eval time = 1404023.25 ms / 32000 tokens ( 43.88 ms per token, 22.79 tokens per second)
392.09.927.225 I slot print_timing: id 0 | task 145539 | total time = 1412661.94 ms / 54269 tokens
392.09.927.225 I slot print_timing: id 0 | task 145539 | graphs reused = 84415
392.09.927.226 I slot print_timing: id 0 | task 145539 | draft acceptance = 0.00000 ( 0 accepted / 127986 generated)
392.09.927.241 I statistics draft-mtp: #calls(b,g,a) = 567 174600 139913, #gen drafts = 139913, #acc drafts = 103668, #gen tokens = 422672, #acc tokens = 274174, dur(b,g,a) = 0.412, 1117416.901, 152.351 ms
392.09.928.237 I slot release: id 0 | task 145539 | stop processing: n_tokens = 67010, truncated = 0
392.09.928.350 I srv update_slots: all slots are idle
392.10.458.829 I srv params_from_: Chat format: peg-native
392.10.460.468 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = 23529366375
392.10.460.471 I srv get_availabl: updating prompt cache
392.10.465.357 W srv prompt_save: - saving prompt with length 67010, total state size = 2168.536 MiB (draft: 119.887 MiB)
392.10.866.297 I srv load: - looking for better prompt, base f_keep = 0.037, sim = 0.044
392.10.866.328 I srv load: - found better prompt with f_keep = 1.000, sim = 0.999
392.11.176.092 I srv update: - cache state: 1 prompts, 3829.656 MiB (limits: 16000.000 MiB, 148224 tokens, 279962 est)
392.11.176.099 I srv update: - prompt 000002142B448BF0: 67010 tokens, checkpoints: 9, 3829.656 MiB
392.11.176.100 I srv get_availabl: prompt cache update took 715.63 ms
392.11.177.093 I reasoning-budget: activated, budget=2147483647 tokens
392.11.177.383 I slot launch_slot_: id 0 | task 177552 | processing task, is_child = 0
392.11.177.427 W slot update_slots: id 0 | task 177552 | old: ...
</tool_call><|im_end|>
392.11.177.428 W slot update_slots: id 0 | task 177552 | new: ...
</tool_call><|im_end|>
392.11.177.428 W slot update_slots: id 0 | task 177552 | 198 248059 248046 198
392.11.177.429 W slot update_slots: id 0 | task 177552 | 198 248059 248046 198
392.11.177.431 W slot update_slots: id 0 | task 177552 | n_past = 56287, slot.prompt.tokens.size() = 56287, seq_id = 0, pos_min = 56286, n_swa = 0
392.11.177.431 I slot update_slots: id 0 | task 177552 | Checking checkpoint with [55702, 55702] against 56286...
392.11.207.188 W slot update_slots: id 0 | task 177552 | restored context checkpoint (pos_min = 55702, pos_max = 55702, n_tokens = 55703, n_past = 55703, size = 249.284 MiB)
392.11.292.435 W slot create_check: id 0 | task 177552 | erasing old context checkpoint (pos_min = 37194, pos_max = 37194, n_tokens = 37195, size = 216.171 MiB)
392.11.348.400 I slot create_check: id 0 | task 177552 | created context checkpoint 32 of 32 (pos_min = 55830, pos_max = 55830, n_tokens = 55831, size = 249.513 MiB)
392.11.610.000 W slot create_check: id 0 | task 177552 | erasing old context checkpoint (pos_min = 37706, pos_max = 37706, n_tokens = 37707, size = 217.087 MiB)
392.11.656.271 I slot create_check: id 0 | task 177552 | created context checkpoint 32 of 32 (pos_min = 56342, pos_max = 56342, n_tokens = 56343, size = 250.429 MiB)
392.16.164.874 I slot print_timing: id 0 | task 177552 | n_decoded = 100, tg = 22.39 t/s
392.19.185.299 I slot print_timing: id 0 | task 177552 | n_decoded = 167, tg = 22.31 t/s
...
Name and Version
.\llama-server.exe --version
version: 196 (40d5358)
built with MSVC 19.44.35226.0 for x64
Operating systems
Windows
GGML backends
CUDA
Hardware
Ryzen 7950X3D + RTX 4090 (driver 596.36)
Models
unsloth/Qwen3.6-27B-MTP-GGUF IQ4_XS
-download validated with sha256
Problem description & steps to reproduce
The server started outputting "///////////////" in a loop, after a very long session (hours), in which it functioned correctly and with good performance.
build with:
cmake -B build -DGGML_NATIVE=ON -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DLLAMA_BUILD_UI=OFF && cmake --build build --config Release -j 32 --target llama-serverrun server:
$env:LLAMA_SERVER_SLOTS_DEBUG=1; llama-server.exe -m unsloth\Qwen3.6-27B-MTP-GGUF\Qwen3.6-27B-IQ4_XS.gguf --host 0.0.0.0 --port 8081 -fa on -ctk q8_0 -ctv q5_1 -ngl 99 -ngld 99 --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0 --repeat-penalty 1 --presence-penalty 0 -c 148000 --parallel 1 --jinja --chat-template-kwargs '{"preserve_thinking": true}' --metrics --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -cram 16000 --log-timestamps --path .\webui\in OpenCode instruct the agent to implement a complicated software project with no intervention (delegate each task to a subagent to conserve context etc.) and let it run
First Bad Commit
No response
Relevant log output
NB: OpenCode has model config output limit at 32k.
Server log snippet (full below) - generation was OK and then it started outputting "////" in a loop until 32k limit
Server init log (verbosity set to 4)
Console copy-paste from Powershell (~19k lines):
llama server MTP loop console log.txt