Skip to content

Eval bug: test-save-load-state aborts on hrm_text with GGML_SCHED_NO_REALLOC (input-layer tensor makes the backend assignment depend on batch size) #29484

Description

@CHIPMUNK-T0T

Name and Version

version: b11194-5-g86a24a182 (master 86a24a1)

Operating systems

Linux

GGML backends

CUDA, HIP, Vulkan

Hardware

RTX 4070 (local).
ggml-ci: T4, ROCm, Intel PTL, Apple.

Models

hrm_text-dense.gguf from test-generate-models (any LLM_ARCH_HRM_TEXT model).

Problem description & steps to reproduce

Test 9 added by #27530 exposes an existing HRM-Text issue under GGML_SCHED_NO_REALLOC: test-save-load-state aborts on hrm_text-dense.gguf in several GPU ctest jobs of ggml-ci (CUDA, ROCm, Vulkan on Intel Linux / NVIDIA / Apple) with unexpected graph reallocation (for example: https://github.com/ggml-org/llama.cpp/actions/runs/36226612703/job/108361769531#step:3:1272). The CPU job passes.

The failure is not in the restore path itself. Test 9 first decodes 8 tokens to obtain the baseline, clears memory, and then decodes 24 tokens. The abort occurs during this second llama_decode() (tests/test-save-load-state.cpp:721), before any corrupted state is restored.

Cause: hrm.z_l_init is registered as LLM_TENSOR_LAYER_INPUT (src/llama-arch.cpp), so it lives in host memory. The ggml_add() with the embeddings in src/models/hrm-text.cpp therefore goes through the scheduler's offload rule and is assigned to the GPU only for batches >= 32 rows. The reserve graph (100 tokens) puts it on the GPU, small decodes put it on the CPU, so the allocator is re-reserved for the 8-token graph and the 24-token decode needs another reallocation (ggml_gallocr_needs_realloc: node embd is not valid), which GGML_SCHED_NO_REALLOC turns into the abort. Without that option it only costs an extra reallocation.

LLM_TENSOR_TOKEN_EMBD_NORM avoids the same problem by being placed on the first layer instead of the input layer; doing the same for hrm.z_l_init (LLM_TENSOR_LAYER_REPEATING, bid 0) makes test-save-load-state --models pass 127/127 on CUDA and CPU with GGML_SCHED_NO_REALLOC=ON.

Repro: build with -DGGML_CUDA=ON -DGGML_SCHED_NO_REALLOC=ON, then test-save-load-state -m tests/test-models/hrm_text-dense.gguf.

First Bad Commit

08618ff (#27530) exposes it by adding test 9; the placement dates from 7d6f5d0 (#27625).

Relevant log output

Logs
hrm_text-dense.gguf      ggml/src/ggml-backend.cpp:1622: ggml_backend_sched_alloc_splits: unexpected graph reallocation (graph size = 385, nodes = 385, leafs = 80), debug_realloc = 1
#10 test_state_restore_failure (...) at tests/test-save-load-state.cpp:721

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions