Skip to content

Misc. bug: Incapable IGP crashes llama.cpp when capable GPU is present #23152

Description

@galmok

Name and Version

C:\Tools\llama.cpp>.\llama-b9174-bin-win-hip-radeon-x64\llama-cli --version
version: 9174 (59778f0)
built with Clang 19.1.5 for Windows x86_64

Operating systems

Windows

Which llama.cpp modules do you know to be affected?

llama-server

Command line

.\llama-b9174-bin-win-hip-radeon-x64\llama-server -m "T:\LLMs\unsloth\Qwen3.6-35B-A3B-GGUF\Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf" --jinja --ctx-size 80000 --no-warmup --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 --presence-penalty 1.5 --reasoning on --mmap -ndio -fitc 80000 -fitt 256 -np 1 -ctk q8_0 -ctv q8_0 -nkvo --port 8088

Problem description & steps to reproduce

There is a regression, seemingly not caused by changes in llama.cpp, but due to a change in AMD GPU drivers on Windows.
With GPU 26.3.1 on AMD 9070 XT (and having IGPU active), cause no problems with llama.cpp until I upgraded the driver. Now, llama.cpp experiences a critical error (crashes; log shown below) during startup. Choice of model doesn't seem to matter.
I have tried working around this by selecting the 9070 XT with -dev ROCm1, but while the behavior changes slightly, llama.cpp still crashes.
I have to use an environment variable to block the knowledge of the IGPU from llama.cpp and then it works.
This is the environment variable required to use llama.cpp now (when the IGPU is number 0):

HIP_VISIBLE_DEVICES=1

This means only show AMD device 1 (which is the 9070 XT in my system) to the application (llama.cpp in this case).

I usually have these defined, but tried unsetting them for this test and they made no difference:

HIP_PATH=C:\Program Files\AMD\ROCm\7.1\
HIP_PATH_71=C:\Program Files\AMD\ROCm\7.1\
ROCBLAS_TENSILE_LIBPATH=C:\Program Files\AMD\ROCm\7.1\bin\rocblas\library

First Bad Commit

No response

Relevant log output

Logs
0.00.039.811 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.039.816 I device_info:
0.00.132.853 I   - ROCm0   : AMD Radeon(TM) Graphics (50246 MiB, 50095 MiB free)
0.00.229.969 I   - ROCm1   : AMD Radeon RX 9070 XT (16304 MiB, 16152 MiB free)
0.00.229.976 I   - CPU     : AMD Ryzen 7 9850X3D 8-Core Processor            (128528 MiB, 100870 MiB free)
0.00.230.028 I system_info: n_threads = 8 (n_threads_batch = 8) / 16 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
0.00.230.063 I srv          init: running without SSL
0.00.230.081 I srv          init: using 15 threads for HTTP server
0.00.230.161 I srv         start: binding port with default address family
0.00.235.001 I srv          main: loading model
0.00.235.022 I srv    load_model: loading model 'T:\LLMs\unsloth\Qwen3.6-35B-A3B-GGUF\Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf'
0.00.235.089 I common_init_result: fitting params to device memory ...
0.00.235.090 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
0.11.125.724 W llama_context: n_ctx_seq (80128) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.11.294.478 I srv    load_model: initializing slots, n_slots = 1
D:/a/llama.cpp/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu:102: ROCm error
0.11.317.358 E ggml_cuda_compute_forward: MUL_MAT failed

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions