Skip to content

Eval bug: MTP with Qwen3.6 27B outputs repeated //// after long session #23577

Description

@aand18

Name and Version

.\llama-server.exe --version
version: 196 (40d5358)
built with MSVC 19.44.35226.0 for x64

Operating systems

Windows

GGML backends

CUDA

Hardware

Ryzen 7950X3D + RTX 4090 (driver 596.36)

Models

unsloth/Qwen3.6-27B-MTP-GGUF IQ4_XS

-download validated with sha256

Problem description & steps to reproduce

The server started outputting "///////////////" in a loop, after a very long session (hours), in which it functioned correctly and with good performance.

  1. build with:
    cmake -B build -DGGML_NATIVE=ON -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DLLAMA_BUILD_UI=OFF && cmake --build build --config Release -j 32 --target llama-server

    nvcc --version
    Cuda compilation tools, release 13.1, V13.1.80
    Build cuda_13.1.r13.1/compiler.36836380_0
    
  2. run server:

    $env:LLAMA_SERVER_SLOTS_DEBUG=1; llama-server.exe -m unsloth\Qwen3.6-27B-MTP-GGUF\Qwen3.6-27B-IQ4_XS.gguf --host 0.0.0.0 --port 8081 -fa on -ctk q8_0 -ctv q5_1 -ngl 99 -ngld 99 --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0 --repeat-penalty 1 --presence-penalty 0 -c 148000 --parallel 1 --jinja --chat-template-kwargs '{"preserve_thinking": true}' --metrics --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -cram 16000 --log-timestamps --path .\webui\

  3. in OpenCode instruct the agent to implement a complicated software project with no intervention (delegate each task to a subagent to conserve context etc.) and let it run

First Bad Commit

No response

Relevant log output

NB: OpenCode has model config output limit at 32k.

Server log snippet (full below) - generation was OK and then it started outputting "////" in a loop until 32k limit
368.34.643.407 W slot update_slots: id  0 | task 145480 |       29     198  248059  248046
368.34.643.407 W slot update_slots: id  0 | task 145480 |       29     198  248059  248046
368.34.643.409 W slot update_slots: id  0 | task 145480 | n_past = 12314, slot.prompt.tokens.size() = 12314, seq_id = 0, pos_min = 12313, n_swa = 0
368.34.643.410 I slot update_slots: id  0 | task 145480 | Checking checkpoint with [12218, 12218] against 12313...
368.34.657.656 W slot update_slots: id  0 | task 145480 | restored context checkpoint (pos_min = 12218, pos_max = 12218, n_tokens = 12219, n_past = 12219, size = 171.487 MiB)
368.34.894.607 I slot create_check: id  0 | task 145480 | created context checkpoint 5 of 32 (pos_min = 12741, pos_max = 12741, n_tokens = 12742, size = 172.423 MiB)
368.35.241.319 I reasoning-budget: deactivated (natural end)
368.35.866.875 I slot print_timing: id  0 | task 145480 | n_decoded =    103, tg = 110.13 t/s
368.37.090.834 I slot print_timing: id  0 | task 145480 | prompt eval time =     288.08 ms /   527 tokens (    0.55 ms per token,  1829.35 tokens per second)
368.37.090.838 I slot print_timing: id  0 | task 145480 |        eval time =    2159.22 ms /   254 tokens (    8.50 ms per token,   117.64 tokens per second)
368.37.090.839 I slot print_timing: id  0 | task 145480 |       total time =    2447.30 ms /   781 tokens
368.37.090.840 I slot print_timing: id  0 | task 145480 |    graphs reused =      52546
368.37.090.841 I slot print_timing: id  0 | task 145480 | draft acceptance = 0.99505 (  201 accepted /   202 generated)
368.37.090.855 I statistics        draft-mtp: #calls(b,g,a) =  566 142602 107915, #gen drafts = 107915, #acc drafts = 103668, #gen tokens = 294686, #acc tokens = 274174, dur(b,g,a) = 0.412, 851457.082, 118.639 ms
368.37.091.409 I slot      release: id  0 | task 145480 | stop processing: n_tokens = 13002, truncated = 0
368.37.091.461 I srv  update_slots: all slots are idle
368.37.263.055 I srv  params_from_: Chat format: peg-native
368.37.264.238 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.371 (> 0.100 thold), f_keep = 1.000
368.37.265.044 I reasoning-budget: activated, budget=2147483647 tokens
368.37.265.250 I slot launch_slot_: id  0 | task 145539 | processing task, is_child = 0
368.37.265.276 W slot update_slots: id  0 | task 145539 | old: ...
</tool_call><|im_end|>
 | <|endoftext|>
368.37.265.276 W slot update_slots: id  0 | task 145539 | new: ...
</tool_call><|im_end|>
 | <|im_start|>
368.37.265.276 W slot update_slots: id  0 | task 145539 |      198  248059  248046     198  248044
368.37.265.277 W slot update_slots: id  0 | task 145539 |      198  248059  248046     198  248045
368.37.265.279 W slot update_slots: id  0 | task 145539 | n_past = 13001, slot.prompt.tokens.size() = 13002, seq_id = 0, pos_min = 13001, n_swa = 0
368.37.265.279 I slot update_slots: id  0 | task 145539 | Checking checkpoint with [12741, 12741] against 13000...
368.37.278.750 W slot update_slots: id  0 | task 145539 | restored context checkpoint (pos_min = 12741, pos_max = 12741, n_tokens = 12742, n_past = 12742, size = 172.423 MiB)
368.40.199.291 I slot update_slots: id  0 | task 145539 | 8192 tokens since last checkpoint at 12742, creating new checkpoint during processing at position 22982
368.40.226.084 I slot create_check: id  0 | task 145539 | created context checkpoint 6 of 32 (pos_min = 20933, pos_max = 20933, n_tokens = 20934, size = 187.079 MiB)
368.40.991.434 I slot print_timing: id  0 | task 145539 | prompt processing, n_tokens =  10240, progress = 0.66, t =   3.73 s / 2748.13 tokens per second
368.41.770.369 I slot print_timing: id  0 | task 145539 | prompt processing, n_tokens =  12288, progress = 0.71, t =   4.51 s / 2727.57 tokens per second
368.42.564.058 I slot print_timing: id  0 | task 145539 | prompt processing, n_tokens =  14336, progress = 0.77, t =   5.30 s / 2705.52 tokens per second
368.43.372.331 I slot print_timing: id  0 | task 145539 | prompt processing, n_tokens =  16384, progress = 0.83, t =   6.11 s / 2682.79 tokens per second
368.43.372.546 I slot update_slots: id  0 | task 145539 | 8192 tokens since last checkpoint at 20934, creating new checkpoint during processing at position 31174
368.43.398.955 I slot create_check: id  0 | task 145539 | created context checkpoint 7 of 32 (pos_min = 29125, pos_max = 29125, n_tokens = 29126, size = 201.735 MiB)
368.44.218.959 I slot print_timing: id  0 | task 145539 | prompt processing, n_tokens =  18432, progress = 0.89, t =   6.95 s / 2650.68 tokens per second
368.45.050.753 I slot print_timing: id  0 | task 145539 | prompt processing, n_tokens =  20480, progress = 0.95, t =   7.79 s / 2630.54 tokens per second
368.45.590.574 I slot print_timing: id  0 | task 145539 | prompt processing, n_tokens =  21753, progress = 0.99, t =   8.33 s / 2612.88 tokens per second
368.45.619.382 I slot create_check: id  0 | task 145539 | created context checkpoint 8 of 32 (pos_min = 34494, pos_max = 34494, n_tokens = 34495, size = 211.341 MiB)
368.45.835.055 I slot print_timing: id  0 | task 145539 | prompt processing, n_tokens =  22265, progress = 1.00, t =   8.57 s / 2598.08 tokens per second
368.45.865.881 I slot create_check: id  0 | task 145539 | created context checkpoint 9 of 32 (pos_min = 35006, pos_max = 35006, n_tokens = 35007, size = 212.257 MiB)
368.49.934.350 I slot print_timing: id  0 | task 145539 | n_decoded =    100, tg =  24.81 t/s
368.52.965.132 I slot print_timing: id  0 | task 145539 | n_decoded =    175, tg =  24.78 t/s
368.55.974.722 I slot print_timing: id  0 | task 145539 | n_decoded =    249, tg =  24.73 t/s

...

392.06.145.199 I slot print_timing: id  0 | task 145539 | n_decoded =  31920, tg =  22.80 t/s
392.09.171.862 I slot print_timing: id  0 | task 145539 | n_decoded =  31984, tg =  22.79 t/s
392.09.927.219 I slot print_timing: id  0 | task 145539 | prompt eval time =    8638.69 ms / 22269 tokens (    0.39 ms per token,  2577.82 tokens per second)
392.09.927.224 I slot print_timing: id  0 | task 145539 |        eval time = 1404023.25 ms / 32000 tokens (   43.88 ms per token,    22.79 tokens per second)
392.09.927.225 I slot print_timing: id  0 | task 145539 |       total time = 1412661.94 ms / 54269 tokens
392.09.927.225 I slot print_timing: id  0 | task 145539 |    graphs reused =      84415
392.09.927.226 I slot print_timing: id  0 | task 145539 | draft acceptance = 0.00000 (    0 accepted / 127986 generated)
392.09.927.241 I statistics        draft-mtp: #calls(b,g,a) =  567 174600 139913, #gen drafts = 139913, #acc drafts = 103668, #gen tokens = 422672, #acc tokens = 274174, dur(b,g,a) = 0.412, 1117416.901, 152.351 ms
392.09.928.237 I slot      release: id  0 | task 145539 | stop processing: n_tokens = 67010, truncated = 0
392.09.928.350 I srv  update_slots: all slots are idle
392.10.458.829 I srv  params_from_: Chat format: peg-native
392.10.460.468 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = 23529366375
392.10.460.471 I srv  get_availabl: updating prompt cache
392.10.465.357 W srv   prompt_save:  - saving prompt with length 67010, total state size = 2168.536 MiB (draft: 119.887 MiB)
392.10.866.297 I srv          load:  - looking for better prompt, base f_keep = 0.037, sim = 0.044
392.10.866.328 I srv          load:  - found better prompt with f_keep = 1.000, sim = 0.999
392.11.176.092 I srv        update:  - cache state: 1 prompts, 3829.656 MiB (limits: 16000.000 MiB, 148224 tokens, 279962 est)
392.11.176.099 I srv        update:    - prompt 000002142B448BF0:   67010 tokens, checkpoints:  9,  3829.656 MiB
392.11.176.100 I srv  get_availabl: prompt cache update took 715.63 ms
392.11.177.093 I reasoning-budget: activated, budget=2147483647 tokens
392.11.177.383 I slot launch_slot_: id  0 | task 177552 | processing task, is_child = 0
392.11.177.427 W slot update_slots: id  0 | task 177552 | old: ...
</tool_call><|im_end|>

392.11.177.428 W slot update_slots: id  0 | task 177552 | new: ...
</tool_call><|im_end|>

392.11.177.428 W slot update_slots: id  0 | task 177552 |      198  248059  248046     198
392.11.177.429 W slot update_slots: id  0 | task 177552 |      198  248059  248046     198
392.11.177.431 W slot update_slots: id  0 | task 177552 | n_past = 56287, slot.prompt.tokens.size() = 56287, seq_id = 0, pos_min = 56286, n_swa = 0
392.11.177.431 I slot update_slots: id  0 | task 177552 | Checking checkpoint with [55702, 55702] against 56286...
392.11.207.188 W slot update_slots: id  0 | task 177552 | restored context checkpoint (pos_min = 55702, pos_max = 55702, n_tokens = 55703, n_past = 55703, size = 249.284 MiB)
392.11.292.435 W slot create_check: id  0 | task 177552 | erasing old context checkpoint (pos_min = 37194, pos_max = 37194, n_tokens = 37195, size = 216.171 MiB)
392.11.348.400 I slot create_check: id  0 | task 177552 | created context checkpoint 32 of 32 (pos_min = 55830, pos_max = 55830, n_tokens = 55831, size = 249.513 MiB)
392.11.610.000 W slot create_check: id  0 | task 177552 | erasing old context checkpoint (pos_min = 37706, pos_max = 37706, n_tokens = 37707, size = 217.087 MiB)
392.11.656.271 I slot create_check: id  0 | task 177552 | created context checkpoint 32 of 32 (pos_min = 56342, pos_max = 56342, n_tokens = 56343, size = 250.429 MiB)
392.16.164.874 I slot print_timing: id  0 | task 177552 | n_decoded =    100, tg =  22.39 t/s
392.19.185.299 I slot print_timing: id  0 | task 177552 | n_decoded =    167, tg =  22.31 t/s

...
Server init log (verbosity set to 4)
0.00.017.243 I common_params_print_info: build 196 (40d5358d3) with MSVC 19.44.35226.0 for x64
0.00.017.246 I log_info: verbosity = 4 (adjust with the `-lv N` CLI arg)
0.00.017.246 I device_info:
0.00.137.842 I   - CUDA0   : NVIDIA GeForce RTX 4090 (24563 MiB, 23036 MiB free)
0.00.137.849 I   - CPU     : AMD Ryzen 9 7950X3D 16-Core Processor           (64627 MiB, 41570 MiB free)
0.00.137.892 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | CUDA : ARCHS = 890 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | AVX512 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
0.00.137.916 I srv          init: running without SSL
0.00.137.930 I srv          init: using 31 threads for HTTP server
0.00.137.931 I srv          init: The UI is disabled
0.00.137.932 I srv          init: Use --ui/--no-ui (or deprecated --webui/--no-webui) to enable/disable
0.00.138.042 I srv         start: binding port with default address family
0.00.151.012 I srv  llama_server: loading model
0.00.151.028 I srv    load_model: loading model '.\.cache\lm-studio\models\unsloth\Qwen3.6-27B-MTP-GGUF\Qwen3.6-27B-IQ4_XS.gguf'
0.00.151.066 I common_init_result: fitting params to device memory ...
0.00.151.066 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
0.00.151.070 I common_params_fit_impl: getting device memory data for initial parameters:
0.00.496.782 I common_memory_breakdown_print: | memory breakdown [MiB] | total    free     self   model   context   compute    unaccounted |
0.00.496.786 I common_memory_breakdown_print: |   - CUDA0 (RTX 4090)   | 24563 = 22286 + (19736 = 14285 +    4945 +     505) +      -17459 |
0.00.496.787 I common_memory_breakdown_print: |   - Host               |                    991 =   682 +       0 +     309                |
0.00.539.742 I common_params_fit_impl: projected to use 19736 MiB of device memory vs. 22286 MiB of free device memory
0.00.539.746 I common_params_fit_impl: will leave 2549 >= 1024 MiB of free device memory, no changes needed
0.00.539.749 I common_fit_params: successfully fit params to free device memory
0.00.539.755 I common_fit_params: fitting params to free memory took 0.25 seconds
0.00.588.519 I llama_model_loader: loaded meta data with 52 key-value pairs and 866 tensors from .\.cache\lm-studio\models\unsloth\Qwen3.6-27B-MTP-GGUF\Qwen3.6-27B-IQ4_XS.gguf (version GGUF V3 (latest))
0.00.588.536 I llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
0.00.588.538 I llama_model_loader: - kv   0:                       general.architecture str              = qwen35
0.00.588.539 I llama_model_loader: - kv   1:                               general.type str              = model
0.00.588.540 I llama_model_loader: - kv   2:                     general.sampling.top_k i32              = 20
0.00.588.544 I llama_model_loader: - kv   3:                     general.sampling.top_p f32              = 0.950000
0.00.588.545 I llama_model_loader: - kv   4:                      general.sampling.temp f32              = 1.000000
0.00.588.545 I llama_model_loader: - kv   5:                               general.name str              = Qwen3.6-27B
0.00.588.546 I llama_model_loader: - kv   6:                           general.basename str              = Qwen3.6-27B
0.00.588.546 I llama_model_loader: - kv   7:                       general.quantized_by str              = Unsloth
0.00.588.546 I llama_model_loader: - kv   8:                         general.size_label str              = 27B
0.00.588.547 I llama_model_loader: - kv   9:                            general.license str              = apache-2.0
0.00.588.549 I llama_model_loader: - kv  10:                       general.license.link str              = https://huggingface.co/Qwen/Qwen3.6-2...
0.00.588.549 I llama_model_loader: - kv  11:                           general.repo_url str              = https://huggingface.co/unsloth
0.00.588.550 I llama_model_loader: - kv  12:                   general.base_model.count u32              = 1
0.00.588.551 I llama_model_loader: - kv  13:                  general.base_model.0.name str              = Qwen3.6 27B
0.00.588.551 I llama_model_loader: - kv  14:          general.base_model.0.organization str              = Qwen
0.00.588.552 I llama_model_loader: - kv  15:              general.base_model.0.repo_url str              = https://huggingface.co/Qwen/Qwen3.6-27B
0.00.588.560 I llama_model_loader: - kv  16:                               general.tags arr[str,2]       = ["unsloth", "image-text-to-text"]
0.00.588.561 I llama_model_loader: - kv  17:                         qwen35.block_count u32              = 65
0.00.588.561 I llama_model_loader: - kv  18:                      qwen35.context_length u32              = 262144
0.00.588.562 I llama_model_loader: - kv  19:                    qwen35.embedding_length u32              = 5120
0.00.588.562 I llama_model_loader: - kv  20:                 qwen35.feed_forward_length u32              = 17408
0.00.588.563 I llama_model_loader: - kv  21:                qwen35.attention.head_count u32              = 24
0.00.588.563 I llama_model_loader: - kv  22:             qwen35.attention.head_count_kv u32              = 4
0.00.588.565 I llama_model_loader: - kv  23:             qwen35.rope.dimension_sections arr[i32,4]       = [11, 11, 10, 0]
0.00.588.566 I llama_model_loader: - kv  24:                      qwen35.rope.freq_base f32              = 10000000.000000
0.00.588.567 I llama_model_loader: - kv  25:    qwen35.attention.layer_norm_rms_epsilon f32              = 0.000001
0.00.588.568 I llama_model_loader: - kv  26:                qwen35.attention.key_length u32              = 256
0.00.588.568 I llama_model_loader: - kv  27:              qwen35.attention.value_length u32              = 256
0.00.588.568 I llama_model_loader: - kv  28:                     qwen35.ssm.conv_kernel u32              = 4
0.00.588.569 I llama_model_loader: - kv  29:                      qwen35.ssm.state_size u32              = 128
0.00.588.569 I llama_model_loader: - kv  30:                     qwen35.ssm.group_count u32              = 16
0.00.588.569 I llama_model_loader: - kv  31:                  qwen35.ssm.time_step_rank u32              = 48
0.00.588.570 I llama_model_loader: - kv  32:                      qwen35.ssm.inner_size u32              = 6144
0.00.588.570 I llama_model_loader: - kv  33:             qwen35.full_attention_interval u32              = 4
0.00.588.571 I llama_model_loader: - kv  34:                qwen35.rope.dimension_count u32              = 64
0.00.588.571 I llama_model_loader: - kv  35:                qwen35.nextn_predict_layers u32              = 1
0.00.588.572 I llama_model_loader: - kv  36:                       tokenizer.ggml.model str              = gpt2
0.00.588.572 I llama_model_loader: - kv  37:                         tokenizer.ggml.pre str              = qwen35
0.00.628.257 I llama_model_loader: - kv  38:                      tokenizer.ggml.tokens arr[str,248320]  = ["!", "\"", "#", "$", "%", "&", "'", ...
0.00.643.531 I llama_model_loader: - kv  39:                  tokenizer.ggml.token_type arr[i32,248320]  = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
0.00.683.186 I llama_model_loader: - kv  40:                      tokenizer.ggml.merges arr[str,247587]  = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",...
0.00.683.190 I llama_model_loader: - kv  41:                tokenizer.ggml.eos_token_id u32              = 248046
0.00.683.191 I llama_model_loader: - kv  42:            tokenizer.ggml.padding_token_id u32              = 248055
0.00.683.191 I llama_model_loader: - kv  43:                tokenizer.ggml.bos_token_id u32              = 248044
0.00.683.192 I llama_model_loader: - kv  44:               tokenizer.ggml.add_bos_token bool             = false
0.00.683.196 I llama_model_loader: - kv  45:                    tokenizer.chat_template str              = {%- set image_count = namespace(value...
0.00.683.196 I llama_model_loader: - kv  46:               general.quantization_version u32              = 2
0.00.683.197 I llama_model_loader: - kv  47:                          general.file_type u32              = 30
0.00.683.198 I llama_model_loader: - kv  48:                      quantize.imatrix.file str              = Qwen3.6-27B-GGUF/imatrix_unsloth.gguf
0.00.683.199 I llama_model_loader: - kv  49:                   quantize.imatrix.dataset str              = unsloth_calibration_Qwen3.6-27B.txt
0.00.683.199 I llama_model_loader: - kv  50:             quantize.imatrix.entries_count u32              = 496
0.00.683.200 I llama_model_loader: - kv  51:              quantize.imatrix.chunks_count u32              = 76
0.00.683.200 I llama_model_loader: - type  f32:  456 tensors
0.00.683.201 I llama_model_loader: - type q8_0:    1 tensors
0.00.683.201 I llama_model_loader: - type q4_K:    7 tensors
0.00.683.201 I llama_model_loader: - type q5_K:  113 tensors
0.00.683.202 I llama_model_loader: - type q6_K:    1 tensors
0.00.683.202 I llama_model_loader: - type iq4_xs:  288 tensors
0.00.683.203 I print_info: file format = GGUF V3 (latest)
0.00.683.204 I print_info: file type   = IQ4_XS - 4.25 bpw
0.00.683.207 I print_info: file size   = 14.62 GiB (4.60 BPW)
0.00.683.241 I llama_prepare_model_devices: using device CUDA0 (NVIDIA GeForce RTX 4090) (0000:01:00.0) - 23036 MiB free
0.00.774.052 I load: 0 unused tokens
0.00.799.350 I load: printing all EOG tokens:
0.00.799.354 I load:   - 248044 ('<|endoftext|>')
0.00.799.354 I load:   - 248046 ('<|im_end|>')
0.00.799.354 I load:   - 248063 ('<|fim_pad|>')
0.00.799.355 I load:   - 248064 ('<|repo_name|>')
0.00.799.355 I load:   - 248065 ('<|file_sep|>')
0.00.799.516 I load: special tokens cache size = 33
0.00.837.085 I load: token to piece cache size = 1.7581 MB
0.00.837.098 I print_info: arch                  = qwen35
0.00.837.099 I print_info: vocab_only            = 0
0.00.837.099 I print_info: no_alloc              = 0
0.00.837.100 I print_info: n_ctx_train           = 262144
0.00.837.100 I print_info: n_embd                = 5120
0.00.837.100 I print_info: n_embd_inp            = 5120
0.00.837.101 I print_info: n_layer               = 65
0.00.837.109 I print_info: n_head                = 24
0.00.837.111 I print_info: n_head_kv             = 4
0.00.837.112 I print_info: n_rot                 = 64
0.00.837.112 I print_info: n_swa                 = 0
0.00.837.112 I print_info: is_swa_any            = 0
0.00.837.113 I print_info: n_embd_head_k         = 256
0.00.837.113 I print_info: n_embd_head_v         = 256
0.00.837.114 I print_info: n_gqa                 = 6
0.00.837.116 I print_info: n_embd_k_gqa          = 1024
0.00.837.117 I print_info: n_embd_v_gqa          = 1024
0.00.837.118 I print_info: f_norm_eps            = 0.0e+00
0.00.837.119 I print_info: f_norm_rms_eps        = 1.0e-06
0.00.837.119 I print_info: f_clamp_kqv           = 0.0e+00
0.00.837.119 I print_info: f_max_alibi_bias      = 0.0e+00
0.00.837.120 I print_info: f_logit_scale         = 0.0e+00
0.00.837.120 I print_info: f_attn_scale          = 0.0e+00
0.00.837.120 I print_info: f_attn_value_scale    = 0.0000
0.00.837.122 I print_info: n_ff                  = 17408
0.00.837.122 I print_info: n_expert              = 0
0.00.837.122 I print_info: n_expert_used         = 0
0.00.837.123 I print_info: n_expert_groups       = 0
0.00.837.123 I print_info: n_group_used          = 0
0.00.837.123 I print_info: causal attn           = 1
0.00.837.124 I print_info: pooling type          = -1
0.00.837.124 I print_info: rope type             = 40
0.00.837.124 I print_info: rope scaling          = linear
0.00.837.125 I print_info: freq_base_train       = 10000000.0
0.00.837.126 I print_info: freq_scale_train      = 1
0.00.837.126 I print_info: n_ctx_orig_yarn       = 262144
0.00.837.126 I print_info: rope_yarn_log_mul     = 0.0000
0.00.837.126 I print_info: rope_finetuned        = unknown
0.00.837.127 I print_info: mrope sections        = [11, 11, 10, 0]
0.00.837.127 I print_info: ssm_d_conv            = 4
0.00.837.127 I print_info: ssm_d_inner           = 6144
0.00.837.128 I print_info: ssm_d_state           = 128
0.00.837.128 I print_info: ssm_dt_rank           = 48
0.00.837.128 I print_info: ssm_n_group           = 16
0.00.837.128 I print_info: ssm_dt_b_c_rms        = 0
0.00.837.129 I print_info: model type            = 27B
0.00.837.129 I print_info: model params          = 27.32 B
0.00.837.130 I print_info: general.name          = Qwen3.6-27B
0.00.837.130 I print_info: vocab type            = BPE
0.00.837.131 I print_info: n_vocab               = 248320
0.00.837.131 I print_info: n_merges              = 247587
0.00.837.131 I print_info: BOS token             = 248044 '<|endoftext|>'
0.00.837.132 I print_info: EOS token             = 248046 '<|im_end|>'
0.00.837.132 I print_info: EOT token             = 248046 '<|im_end|>'
0.00.837.132 I print_info: PAD token             = 248055 '<|vision_pad|>'
0.00.837.133 I print_info: LF token              = 198 'Ċ'
0.00.837.133 I print_info: FIM PRE token         = 248060 '<|fim_prefix|>'
0.00.837.133 I print_info: FIM SUF token         = 248062 '<|fim_suffix|>'
0.00.837.133 I print_info: FIM MID token         = 248061 '<|fim_middle|>'
0.00.837.134 I print_info: FIM PAD token         = 248063 '<|fim_pad|>'
0.00.837.134 I print_info: FIM REP token         = 248064 '<|repo_name|>'
0.00.837.134 I print_info: FIM SEP token         = 248065 '<|file_sep|>'
0.00.837.134 I print_info: EOG token             = 248044 '<|endoftext|>'
0.00.837.135 I print_info: EOG token             = 248046 '<|im_end|>'
0.00.837.135 I print_info: EOG token             = 248063 '<|fim_pad|>'
0.00.837.135 I print_info: EOG token             = 248064 '<|repo_name|>'
0.00.837.136 I print_info: EOG token             = 248065 '<|file_sep|>'
0.00.837.136 I print_info: max token length      = 256
0.00.837.137 I load_tensors: loading model tensors, this can take a while... (mmap = true, direct_io = false)
0.00.943.996 I load_tensors: offloading output layer to GPU
0.00.944.000 I load_tensors: offloading 64 repeating layers to GPU
0.00.944.000 I load_tensors: offloaded 66/66 layers to GPU
0.00.944.006 I load_tensors:   CPU_Mapped model buffer size =   682.03 MiB
0.00.944.007 I load_tensors:        CUDA0 model buffer size = 14285.76 MiB
...........................................................................................
0.04.695.711 I common_init_result: added <|endoftext|> logit bias = -inf
0.04.695.714 I common_init_result: added <|im_end|> logit bias = -inf
0.04.695.715 I common_init_result: added <|fim_pad|> logit bias = -inf
0.04.695.715 I common_init_result: added <|repo_name|> logit bias = -inf
0.04.695.716 I common_init_result: added <|file_sep|> logit bias = -inf
0.04.696.046 I llama_context: constructing llama_context
0.04.696.049 I llama_context: n_seq_max     = 1
0.04.696.049 I llama_context: n_ctx         = 148224
0.04.696.049 I llama_context: n_ctx_seq     = 148224
0.04.696.050 I llama_context: n_batch       = 2048
0.04.696.050 I llama_context: n_ubatch      = 512
0.04.696.050 I llama_context: causal_attn   = 1
0.04.696.050 I llama_context: flash_attn    = enabled
0.04.696.051 I llama_context: kv_unified    = false
0.04.696.053 I llama_context: freq_base     = 10000000.0
0.04.696.054 I llama_context: freq_scale    = 1
0.04.696.054 I llama_context: n_rs_seq      = 4
0.04.696.054 W llama_context: n_ctx_seq (148224) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.04.696.323 I llama_context:  CUDA_Host  output buffer size =     0.95 MiB
0.04.703.360 I llama_kv_cache:      CUDA0 KV buffer size =  4197.75 MiB
0.04.791.920 I llama_kv_cache: size = 4197.75 MiB (148224 cells,  16 layers,  1/1 seqs), K (q8_0): 2460.75 MiB, V (q5_1): 1737.00 MiB
0.04.791.928 I llama_kv_cache: attn_rot_k = 1, n_embd_head_k_all = 256
0.04.791.928 I llama_kv_cache: attn_rot_v = 1, n_embd_head_k_all = 256
0.04.809.433 I llama_memory_recurrent:      CUDA0 RS buffer size =   748.12 MiB
0.04.809.442 I llama_memory_recurrent: size =  748.12 MiB (     1 cells,  65 layers,  1 seqs  4 rs_seq), R (f32):   28.12 MiB, S (f32):  720.00 MiB
0.04.809.447 I sched_reserve: reserving ...
0.04.811.569 I sched_reserve: resolving fused Gated Delta Net support:
0.04.812.739 I sched_reserve: fused Gated Delta Net (autoregressive) enabled
0.04.813.639 I sched_reserve: fused Gated Delta Net (chunked) enabled
0.04.836.151 I sched_reserve:      CUDA0 compute buffer size =   505.00 MiB
0.04.836.156 I sched_reserve:  CUDA_Host compute buffer size =   309.79 MiB
0.04.836.157 I sched_reserve: graph nodes  = 5096
0.04.836.157 I sched_reserve: graph splits = 2
0.04.836.158 I sched_reserve: reserve took 26.71 ms, sched copies = 1
0.04.836.421 I common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
0.04.937.146 I srv    load_model: creating MTP draft context against the target model '.\.cache\lm-studio\models\unsloth\Qwen3.6-27B-MTP-GGUF\Qwen3.6-27B-IQ4_XS.gguf'
0.04.937.186 I llama_context: constructing llama_context
0.04.937.189 I llama_context: n_seq_max     = 1
0.04.937.189 I llama_context: n_ctx         = 148224
0.04.937.189 I llama_context: n_ctx_seq     = 148224
0.04.937.189 I llama_context: n_batch       = 2048
0.04.937.190 I llama_context: n_ubatch      = 512
0.04.937.190 I llama_context: causal_attn   = 1
0.04.937.190 I llama_context: flash_attn    = enabled
0.04.937.191 I llama_context: kv_unified    = false
0.04.937.193 I llama_context: freq_base     = 10000000.0
0.04.937.194 I llama_context: freq_scale    = 1
0.04.937.194 I llama_context: n_rs_seq      = 0
0.04.937.194 W llama_context: n_ctx_seq (148224) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.04.937.241 I llama_context:  CUDA_Host  output buffer size =     0.95 MiB
0.04.938.679 I llama_kv_cache:      CUDA0 KV buffer size =   262.36 MiB
0.04.944.075 I llama_kv_cache: size =  262.36 MiB (148224 cells,   1 layers,  1/1 seqs), K (q8_0):  153.80 MiB, V (q5_1):  108.56 MiB
0.04.944.081 I llama_kv_cache: attn_rot_k = 1, n_embd_head_k_all = 256
0.04.944.081 I llama_kv_cache: attn_rot_v = 1, n_embd_head_k_all = 256
0.04.944.189 I sched_reserve: reserving ...
0.04.946.187 I sched_reserve: resolving fused Gated Delta Net support:
0.04.946.301 I sched_reserve: fused Gated Delta Net (autoregressive) enabled
0.04.946.354 I sched_reserve: fused Gated Delta Net (chunked) enabled
0.04.963.635 I sched_reserve:      CUDA0 compute buffer size =   505.00 MiB
0.04.963.640 I sched_reserve:  CUDA_Host compute buffer size =   309.79 MiB
0.04.963.641 I sched_reserve: graph nodes  = 63
0.04.963.641 I sched_reserve: graph splits = 2
0.04.963.641 I sched_reserve: reserve took 19.45 ms, sched copies = 1
0.04.973.061 I srv    load_model: initializing slots, n_slots = 1
0.05.000.968 I common_context_can_seq_rm: the context supports bounded partial sequence removal
0.05.031.564 I common_speculative_impl_draft_mtp: adding speculative implementation 'draft-mtp'
0.05.031.572 I common_speculative_impl_draft_mtp: - n_max=4, n_min=0, p_min=0.75, n_embd=5120, backend_sampling=1
0.05.031.575 I common_speculative_impl_draft_mtp: - gpu_layers=99, cache_k=f16, cache_v=f16, ctx_tgt=yes, ctx_dft=yes, devices=[default]
0.05.031.671 I srv    load_model: speculative decoding context initialized
0.05.031.672 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 148224
0.05.031.683 W srv    load_model: LLAMA_SERVER_SLOTS_DEBUG = 1
0.05.031.732 I srv    load_model: prompt cache is enabled, size limit: 16000 MiB
0.05.031.732 I srv    load_model: use `--cache-ram 0` to disable the prompt cache
0.05.031.732 I srv    load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
0.05.031.748 W srv          init: --cache-idle-slots requires --kv-unified, disabling
0.05.048.908 I init: chat template, example_format: '<|im_start|>system
You are a helpful assistant<|im_end|>
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
<think>

</think>

Hi there<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant
<think>
'
0.05.061.645 I srv          init: init: chat template, thinking = 1
0.05.061.658 I srv  llama_server: model loaded
0.05.061.660 I srv  llama_server: server is listening on http://0.0.0.0:8081
0.05.061.671 I srv  update_slots: all slots are idle
0.08.535.011 I srv   operator (): operator (): cleaning up before exit...
0.08.536.236 I common_memory_breakdown_print: | memory breakdown [MiB] | total   free     self   model   context   compute    unaccounted |
0.08.536.239 I common_memory_breakdown_print: |   - CUDA0 (RTX 4090)   | 24563 = 2065 + (19736 = 14285 +    4945 +     505) +        2760 |
0.08.536.239 I common_memory_breakdown_print: |   - Host               |                   991 =   682 +       0 +     309                |

Console copy-paste from Powershell (~19k lines):

llama server MTP loop console log.txt

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions