Name and Version
version: b11194-5-g86a24a182 (master 86a24a1)
Operating systems
Linux
GGML backends
CUDA, HIP, Vulkan
Hardware
RTX 4070 (local).
ggml-ci: T4, ROCm, Intel PTL, Apple.
Models
hrm_text-dense.gguf from test-generate-models (any LLM_ARCH_HRM_TEXT model).
Problem description & steps to reproduce
Test 9 added by #27530 exposes an existing HRM-Text issue under GGML_SCHED_NO_REALLOC: test-save-load-state aborts on hrm_text-dense.gguf in several GPU ctest jobs of ggml-ci (CUDA, ROCm, Vulkan on Intel Linux / NVIDIA / Apple) with unexpected graph reallocation (for example: https://github.com/ggml-org/llama.cpp/actions/runs/36226612703/job/108361769531#step:3:1272). The CPU job passes.
The failure is not in the restore path itself. Test 9 first decodes 8 tokens to obtain the baseline, clears memory, and then decodes 24 tokens. The abort occurs during this second llama_decode() (tests/test-save-load-state.cpp:721), before any corrupted state is restored.
Cause: hrm.z_l_init is registered as LLM_TENSOR_LAYER_INPUT (src/llama-arch.cpp), so it lives in host memory. The ggml_add() with the embeddings in src/models/hrm-text.cpp therefore goes through the scheduler's offload rule and is assigned to the GPU only for batches >= 32 rows. The reserve graph (100 tokens) puts it on the GPU, small decodes put it on the CPU, so the allocator is re-reserved for the 8-token graph and the 24-token decode needs another reallocation (ggml_gallocr_needs_realloc: node embd is not valid), which GGML_SCHED_NO_REALLOC turns into the abort. Without that option it only costs an extra reallocation.
LLM_TENSOR_TOKEN_EMBD_NORM avoids the same problem by being placed on the first layer instead of the input layer; doing the same for hrm.z_l_init (LLM_TENSOR_LAYER_REPEATING, bid 0) makes test-save-load-state --models pass 127/127 on CUDA and CPU with GGML_SCHED_NO_REALLOC=ON.
Repro: build with -DGGML_CUDA=ON -DGGML_SCHED_NO_REALLOC=ON, then test-save-load-state -m tests/test-models/hrm_text-dense.gguf.
First Bad Commit
08618ff (#27530) exposes it by adding test 9; the placement dates from 7d6f5d0 (#27625).
Relevant log output
Logs
hrm_text-dense.gguf ggml/src/ggml-backend.cpp:1622: ggml_backend_sched_alloc_splits: unexpected graph reallocation (graph size = 385, nodes = 385, leafs = 80), debug_realloc = 1
#10 test_state_restore_failure (...) at tests/test-save-load-state.cpp:721
Name and Version
version: b11194-5-g86a24a182 (master 86a24a1)
Operating systems
Linux
GGML backends
CUDA, HIP, Vulkan
Hardware
RTX 4070 (local).
ggml-ci: T4, ROCm, Intel PTL, Apple.
Models
hrm_text-dense.gguffromtest-generate-models(anyLLM_ARCH_HRM_TEXTmodel).Problem description & steps to reproduce
Test 9 added by #27530 exposes an existing HRM-Text issue under
GGML_SCHED_NO_REALLOC:test-save-load-stateaborts onhrm_text-dense.ggufin several GPU ctest jobs of ggml-ci (CUDA, ROCm, Vulkan on Intel Linux / NVIDIA / Apple) withunexpected graph reallocation(for example: https://github.com/ggml-org/llama.cpp/actions/runs/36226612703/job/108361769531#step:3:1272). The CPU job passes.The failure is not in the restore path itself. Test 9 first decodes 8 tokens to obtain the baseline, clears memory, and then decodes 24 tokens. The abort occurs during this second
llama_decode()(tests/test-save-load-state.cpp:721), before any corrupted state is restored.Cause:
hrm.z_l_initis registered asLLM_TENSOR_LAYER_INPUT(src/llama-arch.cpp), so it lives in host memory. Theggml_add()with the embeddings insrc/models/hrm-text.cpptherefore goes through the scheduler's offload rule and is assigned to the GPU only for batches >= 32 rows. The reserve graph (100 tokens) puts it on the GPU, small decodes put it on the CPU, so the allocator is re-reserved for the 8-token graph and the 24-token decode needs another reallocation (ggml_gallocr_needs_realloc: node embd is not valid), whichGGML_SCHED_NO_REALLOCturns into the abort. Without that option it only costs an extra reallocation.LLM_TENSOR_TOKEN_EMBD_NORMavoids the same problem by being placed on the first layer instead of the input layer; doing the same forhrm.z_l_init(LLM_TENSOR_LAYER_REPEATING, bid 0) makestest-save-load-state --modelspass 127/127 on CUDA and CPU withGGML_SCHED_NO_REALLOC=ON.Repro: build with
-DGGML_CUDA=ON -DGGML_SCHED_NO_REALLOC=ON, thentest-save-load-state -m tests/test-models/hrm_text-dense.gguf.First Bad Commit
08618ff (#27530) exposes it by adding test 9; the placement dates from 7d6f5d0 (#27625).
Relevant log output
Logs