Repository navigation
Conversation
Tested this on the hardware and model that #21576 was opened for. TL;DR: the assert is gone and the saved state is complete again Setup
37-token prompt into slot 0, then Results
Both aborts are * that row is from my earlier run reported in #21576, not rebuilt today; every other row comes from this one checkout. The byte count: 8,337,184 is what the non-splitting path writes. This PR reproduces that number exactly in both splitting configurations, so whether the allocation splits no longer changes what ends up in the state. That is the invariant that was broken. The guards in #21576 write 151,576 bytes less. That is exactly one layer's K plus V for 37 cells (2 × 37 × 1024 × 2 B, plus 24 B of header): the last layer is silently dropped, which is what @liminfei-amd predicted from fault injection. So this PR does not merely remove the assert, it restores the data the guards were throwing away. Restore round-trips in every passing case:
On the
|
|
Follow-up on local A/B testing after an AMD driver update cleared the earlier Same command, 10 runs each on Vulkan:
I also updated the PR Test plan accordingly. Thanks again to @ehotting for the Gemma / slot save-restore validation. |
|
Additional verification for both #23737 and #19839 on the same hardware. Setup
1) #23737 (speculative MTP /
|
| build | BLOCK_SIZE |
result |
|---|---|---|
| before | default (1 GiB) | ASSERT |
| before | 16 GiB (avoid split) | OK |
| before | 256 MiB (force split) | ASSERT |
| after | 256 MiB (force split) | OK |
2) #19839 (server multi-slot / prompt-cache path)
does not fit my machine spec: Vulkan ErrorOutOfDeviceMemory during context init.
Minimal repro that fits:
GGML_VK_SUBALLOCATION_BLOCK_SIZE=268435456 llama-server -m gpt-oss-120b-mxfp4-00001-of-00003.gguf
-c 1048576 --parallel 4 -ngl 999 --no-mmap
-sps 0 --cache-prompt --cache-idle-slots --cache-ram 8192
--host 127.0.0.1 --port 9091
Then sequential /v1/chat/completions with distinct prompts so LRU rotates slots (-sps 0).
| build | result |
|---|---|
| before | req 1 OK; on req 2 (slot 3 -> 2) ASSERT |
| after | reqs 1–6 OK |
Takeaway
Both issues hit the same assert class: host state IO over KV stream views left without ggml_backend_view_init after a buft max_size split. #23737 enters via speculative MTP checkpointing; #19839 via server multi-slot / idle-slot cache save. This PR fixes that missing view init in the allocator.
|
Additional verification for #21762 on the same hardware. Setup
#21762 (server prompt-cache / 2nd request)Then two independent
Crash is With TakeawaySame missing Fixes #21762 |
Confirmed on b10352 / RADV gfx1151 — this PR is the right fixAdopted this patch locally on Hardware: AMD Ryzen AI Max+ 395 / Radeon 8060S (gfx1151), RADV, Mesa 25.2.8, Vulkan. Unified memory, GTT pool 124 GiB. Model: Qwen3.8-27B UD-Q8_K_XL + F16 mmproj, arch Reproduce (stock b10352, no patch) — llama-server onlyNo extra proxy in front. The abort is in Same flags with
Slot 0 can answer. The next host-copy of slot state kills the process in ~1 s (not DeviceLost, not the 60 s amdgpu watchdog):
This is not OOM. 4×204800 allocates and idles at 86.5 / 124 GiB. RADV Why
|
|
The PR fixed the
Original failure trace: |
3f82bd2 to
0074731
Compare
|
Rebased after #27644 updated the base files. Still waiting for review. Thanks. |
Root cause and fix for a separate --parallel>1 draft-mtp SIGSEGV/abort crash: GGML_ASSERT(tensor->data != NULL && "tensor not allocated") aborting the whole process on periodic checkpoint creation (common_prompt_checkpoint::update_dft -> server_context_impl::create_checkpoint -> llama_state_seq_get_data_ext -> ggml_backend_tensor_get), reproduced live on Qwen3.8-27B-UD-Q4_K_XL (Vulkan/gfx1151) with --parallel 3 and draft-mtp speculative decoding, via direct kubectl port-forward + log capture. A live diagnostic build confirmed the crashing tensor is a per-stream KV-cache view (layer 64 MTP k_stream) with data=(nil) buffer=(nil) view_src=<valid, allocated> -- i.e. the VIEW itself was never initialized even though its backing storage was. Note: this is a distinct bug from the empty-draft-sequence crash fixed by the preceding commit -- that one is a genuinely never-touched sequence being read at all; this one is a sequence whose backing view was never initialized even though it had been used and has real data. This is a confirmed, independently-reported-upstream bug, not something specific to this fleet's config: - Issue ggml-org/llama.cpp#23737 ("Eval bug: GGML_ASSERT(tensor->data != NULL) on Vulkan since b9318") -- exact assert match, multiple independent reporters on different hardware (RADEON AI PRO r9700) and different models (gemma-4-31B-it, gemma-4-26B-A4B), confirming this is a general Vulkan-backend multi-stream/view-allocation bug, not MTP- or model-specific. - PR ggml-org/llama.cpp#27738 (closed, not merged, by its own author in favor of the PR below) describes the exact same reproduction: model Qwen3.8-27B-UD-Q4_K_XL, MTP/nextn layer 64, same call stack (common_prompt_checkpoint::update_dft -> server_context_impl::create_checkpoint -> pre_decode). Root cause per that PR's own body: "The k_stream/v_stream views are created before the k/v tensors are allocated. On allocation paths that skip ggml_backend_view_init ... they keep tensor->data == NULL forever while their storage lives on view_src." - PR ggml-org/llama.cpp#25584 ("ggml: fix view init skipped after buft max_size split", open/unmerged as of vendoring) is the fix applied here -- named by #27738's own closing comment as "the better fix," fixing it at the ggml allocator level rather than the KV-cache call site, predating #27738 by six weeks, and additionally covering two related issues (#19839, #21762). ROOT CAUSE: when ggml_backend_alloc_ctx_tensors_from_buft splits a context's allocation across multiple buffers because a tensor would exceed the backend buft max_size (Vulkan's default is 1 GiB -- exactly the class of split a ~27B model's KV-cache tensors can trigger), a view-only tail at the end of a split range can skip alloc_tensor_range entirely, so any persistent view landing in that tail (this fleet's layer-64 MTP k_stream/v_stream) never gets ggml_backend_view_init and keeps data == NULL indefinitely, even though its view_src is fully allocated. FIX: instead of trying to allocate-or-view-init each tensor inline during a single pass over one split range (the old alloc_tensor_range logic, exactly what a view-only tail could skip), allocate ONLY real (non-view) tensors during the per-range pass, then run one final pass over the WHOLE context after all splits complete, initializing any tensor whose view_src is set but whose buffer is still NULL. This guarantees every view gets initialized regardless of which split range its parent tensor's allocation landed in. TEST FILE NOT PORTED: this PR's own tests/test-alloc.cpp addition (test_view_init_after_max_size_split) does not apply cleanly against this commit (context drift from the PR's own last rebase point, same class of drift as patch 0001's grammar fix) and was not hand-re-ported here. Low priority to re-port since the fleet's own build disables LLAMA_BUILD_TESTS anyway; worth doing if this repo's tests are ever turned back on. VERIFIED: applies clean (git apply --check) against this exact pinned commit with zero context drift, and gcc -fsyntax-only clean on the modified file. Live-verified end-to-end on real fleet hardware: the exact original crash repro (>16K-token prompt, --parallel 3, draft-mtp, kv_unified=false, mid-prefill checkpoint) completed cleanly (HTTP 200, 0 pod restarts) against an image built from this commit, where the prior image reliably aborted on an identical repro. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017R3QNP4GznFWhhf11QY18j
Related diagnosis: ggml-org/llama.cpp#25584 Assisted-by: Codex
Root cause and fix for a separate --parallel>1 draft-mtp SIGSEGV/abort crash: GGML_ASSERT(tensor->data != NULL && "tensor not allocated") aborting the whole process on periodic checkpoint creation (common_prompt_checkpoint::update_dft -> server_context_impl::create_checkpoint -> llama_state_seq_get_data_ext -> ggml_backend_tensor_get), reproduced live on Qwen3.8-27B-UD-Q4_K_XL (Vulkan/gfx1151) with --parallel 3 and draft-mtp speculative decoding, via direct kubectl port-forward + log capture. A live diagnostic build confirmed the crashing tensor is a per-stream KV-cache view (layer 64 MTP k_stream) with data=(nil) buffer=(nil) view_src=<valid, allocated> -- i.e. the VIEW itself was never initialized even though its backing storage was. Note: this is a distinct bug from the empty-draft-sequence crash fixed by the preceding commit -- that one is a genuinely never-touched sequence being read at all; this one is a sequence whose backing view was never initialized even though it had been used and has real data. This is a confirmed, independently-reported-upstream bug, not something specific to this fleet's config: - Issue ggml-org/llama.cpp#23737 ("Eval bug: GGML_ASSERT(tensor->data != NULL) on Vulkan since b9318") -- exact assert match, multiple independent reporters on different hardware (RADEON AI PRO r9700) and different models (gemma-4-31B-it, gemma-4-26B-A4B), confirming this is a general Vulkan-backend multi-stream/view-allocation bug, not MTP- or model-specific. - PR ggml-org/llama.cpp#27738 (closed, not merged, by its own author in favor of the PR below) describes the exact same reproduction: model Qwen3.8-27B-UD-Q4_K_XL, MTP/nextn layer 64, same call stack (common_prompt_checkpoint::update_dft -> server_context_impl::create_checkpoint -> pre_decode). Root cause per that PR's own body: "The k_stream/v_stream views are created before the k/v tensors are allocated. On allocation paths that skip ggml_backend_view_init ... they keep tensor->data == NULL forever while their storage lives on view_src." - PR ggml-org/llama.cpp#25584 ("ggml: fix view init skipped after buft max_size split", open/unmerged as of vendoring) is the fix applied here -- named by #27738's own closing comment as "the better fix," fixing it at the ggml allocator level rather than the KV-cache call site, predating #27738 by six weeks, and additionally covering two related issues (#19839, #21762). ROOT CAUSE: when ggml_backend_alloc_ctx_tensors_from_buft splits a context's allocation across multiple buffers because a tensor would exceed the backend buft max_size (Vulkan's default is 1 GiB -- exactly the class of split a ~27B model's KV-cache tensors can trigger), a view-only tail at the end of a split range can skip alloc_tensor_range entirely, so any persistent view landing in that tail (this fleet's layer-64 MTP k_stream/v_stream) never gets ggml_backend_view_init and keeps data == NULL indefinitely, even though its view_src is fully allocated. FIX: instead of trying to allocate-or-view-init each tensor inline during a single pass over one split range (the old alloc_tensor_range logic, exactly what a view-only tail could skip), allocate ONLY real (non-view) tensors during the per-range pass, then run one final pass over the WHOLE context after all splits complete, initializing any tensor whose view_src is set but whose buffer is still NULL. This guarantees every view gets initialized regardless of which split range its parent tensor's allocation landed in. TEST FILE NOT PORTED: this PR's own tests/test-alloc.cpp addition (test_view_init_after_max_size_split) does not apply cleanly against this commit (context drift from the PR's own last rebase point, same class of drift as patch 0001's grammar fix) and was not hand-re-ported here. Low priority to re-port since the fleet's own build disables LLAMA_BUILD_TESTS anyway; worth doing if this repo's tests are ever turned back on. VERIFIED: applies clean (git apply --check) against this exact pinned commit with zero context drift, and gcc -fsyntax-only clean on the modified file. Live-verified end-to-end on real fleet hardware: the exact original crash repro (>16K-token prompt, --parallel 3, draft-mtp, kv_unified=false, mid-prefill checkpoint) completed cleanly (HTTP 200, 0 pod restarts) against an image built from this commit, where the prior image reliably aborted on an identical repro. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017R3QNP4GznFWhhf11QY18j
Root cause and fix for a separate --parallel>1 draft-mtp SIGSEGV/abort crash: GGML_ASSERT(tensor->data != NULL && "tensor not allocated") aborting the whole process on periodic checkpoint creation (common_prompt_checkpoint::update_dft -> server_context_impl::create_checkpoint -> llama_state_seq_get_data_ext -> ggml_backend_tensor_get), reproduced live on Qwen3.8-27B-UD-Q4_K_XL (Vulkan/gfx1151) with --parallel 3 and draft-mtp speculative decoding, via direct kubectl port-forward + log capture. A live diagnostic build confirmed the crashing tensor is a per-stream KV-cache view (layer 64 MTP k_stream) with data=(nil) buffer=(nil) view_src=<valid, allocated> -- i.e. the VIEW itself was never initialized even though its backing storage was. Note: this is a distinct bug from the empty-draft-sequence crash fixed by the preceding commit -- that one is a genuinely never-touched sequence being read at all; this one is a sequence whose backing view was never initialized even though it had been used and has real data. This is a confirmed, independently-reported-upstream bug, not something specific to this fleet's config: - Issue ggml-org/llama.cpp#23737 ("Eval bug: GGML_ASSERT(tensor->data != NULL) on Vulkan since b9318") -- exact assert match, multiple independent reporters on different hardware (RADEON AI PRO r9700) and different models (gemma-4-31B-it, gemma-4-26B-A4B), confirming this is a general Vulkan-backend multi-stream/view-allocation bug, not MTP- or model-specific. - PR ggml-org/llama.cpp#27738 (closed, not merged, by its own author in favor of the PR below) describes the exact same reproduction: model Qwen3.8-27B-UD-Q4_K_XL, MTP/nextn layer 64, same call stack (common_prompt_checkpoint::update_dft -> server_context_impl::create_checkpoint -> pre_decode). Root cause per that PR's own body: "The k_stream/v_stream views are created before the k/v tensors are allocated. On allocation paths that skip ggml_backend_view_init ... they keep tensor->data == NULL forever while their storage lives on view_src." - PR ggml-org/llama.cpp#25584 ("ggml: fix view init skipped after buft max_size split", open/unmerged as of vendoring) is the fix applied here -- named by #27738's own closing comment as "the better fix," fixing it at the ggml allocator level rather than the KV-cache call site, predating #27738 by six weeks, and additionally covering two related issues (#19839, #21762). ROOT CAUSE: when ggml_backend_alloc_ctx_tensors_from_buft splits a context's allocation across multiple buffers because a tensor would exceed the backend buft max_size (Vulkan's default is 1 GiB -- exactly the class of split a ~27B model's KV-cache tensors can trigger), a view-only tail at the end of a split range can skip alloc_tensor_range entirely, so any persistent view landing in that tail (this fleet's layer-64 MTP k_stream/v_stream) never gets ggml_backend_view_init and keeps data == NULL indefinitely, even though its view_src is fully allocated. FIX: instead of trying to allocate-or-view-init each tensor inline during a single pass over one split range (the old alloc_tensor_range logic, exactly what a view-only tail could skip), allocate ONLY real (non-view) tensors during the per-range pass, then run one final pass over the WHOLE context after all splits complete, initializing any tensor whose view_src is set but whose buffer is still NULL. This guarantees every view gets initialized regardless of which split range its parent tensor's allocation landed in. TEST FILE NOT PORTED: this PR's own tests/test-alloc.cpp addition (test_view_init_after_max_size_split) does not apply cleanly against this commit (context drift from the PR's own last rebase point, same class of drift as patch 0001's grammar fix) and was not hand-re-ported here. Low priority to re-port since the fleet's own build disables LLAMA_BUILD_TESTS anyway; worth doing if this repo's tests are ever turned back on. VERIFIED: applies clean (git apply --check) against this exact pinned commit with zero context drift, and gcc -fsyntax-only clean on the modified file. Live-verified end-to-end on real fleet hardware: the exact original crash repro (>16K-token prompt, --parallel 3, draft-mtp, kv_unified=false, mid-prefill checkpoint) completed cleanly (HTTP 200, 0 pod restarts) against an image built from this commit, where the prior image reliably aborted on an identical repro. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017R3QNP4GznFWhhf11QY18j
|
I can provide even more verification, just merged this locally and I'm no longer getting the failed assert with multiple agents... My hardware consists of an R9700, W6800 and MI50 on a 5950x. I'm running Vulkan under Gentoo. I see this with Qwen3.8 and Qwen3.8-Flash-Next. Enabling/disabling MTP doesn't seem to have an impact. Could we get this reviewed/merged? Happy to get involved if I can help in any way. |
ggml_backend_alloc_ctx_tensors_from_buft splits a context across several buffers when a tensor would exceed the buffer type's max_size (1 GiB on Vulkan). Views were initialized inside each split range, so a view whose range was never allocated (a view-only tail after the last parent, or a view of a parent that landed in a later range) kept data == NULL while its view_src was allocated. The KV cache's persistent k_stream/v_stream views are such tensors, and every host state read of the cache (prompt checkpoints, the server's RAM prompt cache, slot save) then asserts "tensor not allocated". A 1M-cell KV pool on this model puts each layer's K and V at ~1 GiB, which is what triggers the split. Allocate only parent tensors per range and initialize every view in one pass after all splits. From ggml-org/llama.cpp PR ggml-org#25584 (Yoshi4470).
When ggml_backend_alloc_ctx_tensors_from_buft splits allocation on buft max_size, a view-only tail at the end of the context could skip the final alloc_tensor_range. Persistent views (e.g. KV k_stream / v_stream) were then left without ggml_backend_view_init. Allocate only parent tensors in alloc_tensor_range and initialize all views in a final pass over the context after all splits complete.
Cover the buft max_size split path where a view-only tail would skip ggml_backend_view_init without the finalize pass.
0074731 to
d8aa674
Compare
|
Rebased onto the latest master after a conflict in the test code. |
When ggml_backend_alloc_ctx_tensors_from_buft splits allocation on buft max_size, a view-only tail at the end of the context could skip the final alloc_tensor_range. Persistent views (e.g. KV k_stream / v_stream) were then left without ggml_backend_view_init.
Allocate only parent tensors in alloc_tensor_range and initialize all views in a final pass over the context after all splits complete.
Overview
ggml_backend_alloc_ctx_tensors_from_buftcan split a context across multiple buffers when a tensor would exceed the backend buffer-typemax_size(for example Vulkan's default 1 GiB). If the remaining tail of the context contains only views, the finalalloc_tensor_rangemay be skipped, so those views never getggml_backend_view_init.Persistent KV stream views (
layer.k_stream/v_stream) then keeptensor->data == NULLwhileview_srcis allocated. Host state save/restore later callsggml_backend_tensor_get/seton those views and hits:GGML_ASSERT(tensor->data != NULL && "tensor not allocated")This change:
Also adds
test_view_init_after_max_size_splitintests/test-alloc.cppto cover the oversized-parent + view-only-tail case.
Additional information
Fixes #19839
Fixes #23737
Fixes #21762
Related:
!tensor->dataguards; complementary discussion)I reproduced the assert with Qwen3.6-27B (MTP) when a parent KV tensor exceeded buft
max_sizeand the context ended with a view-only tail.Note on kv-cache : fix crash in state save/restore #21576: skipping tensors with
!tensor->dataavoids the assert, but can omit real KV still reachable viaview_srcand misalign the serialized state. This PR fixes the missing view init in the allocator instead.Prior investigation commit (same patch on older master; kept for reference):
Yoshi4470@99655d5
Test plan
test-alloc(includes newtest_view_init_after_max_size_split)e3546c794): 0/10 (assert every time)amdvlk64.dllcrash after the assert site was resolved by an AMD driver update; treated as separate from this allocator fixCommand used for the A/B runs:
Requirements