Name and Version
lemonade server 11.7.0 + rocm b10472
Operating systems
Windows
GGML backends
HIP
Hardware
AMD AI MAX 395+ 64 GB (gfx1151)
Models
Qwen3.6:27b and 35b-3ba, Qwen3.8:27b
Problem description & steps to reproduce
vscode + cline 4.1.x + lemonade server + rocm
Current status:
- Tool calling works partially with Qwen3.8:27b:
native Cline tools (read_files, search_codebase, editor, etc.) are correctly
recognized and invoked.
- MCP tools still fail: when MCP servers are connected in Cline and their tools
are injected into the context, the model does not correctly invoke them on the
ROCm backend. The same setup works correctly on the Vulkan backend.
model is loaded on both backends with
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --no-mmap --temp 0.1 --repeat-penalty 1.0 --top-k 20 --top-p 0.95 --min-p 0 --chat-template-kwargs '{"preserve_thinking":true,"reasoning_effort":"low"}'
Hypothesis: the issue may be context-size related — MCP tool schemas significantly
increase the prompt token count, and there may be a ROCm-specific behavior difference
under high context load compared to Vulkan.
First Bad Commit
No response
Relevant log output
I don't have any, let me know if lemonade server logs are enough, I can provide them.
Name and Version
lemonade server 11.7.0 + rocm b10472
Operating systems
Windows
GGML backends
HIP
Hardware
AMD AI MAX 395+ 64 GB (gfx1151)
Models
Qwen3.6:27b and 35b-3ba, Qwen3.8:27b
Problem description & steps to reproduce
vscode + cline 4.1.x + lemonade server + rocm
Current status:
native Cline tools (read_files, search_codebase, editor, etc.) are correctly
recognized and invoked.
are injected into the context, the model does not correctly invoke them on the
ROCm backend. The same setup works correctly on the Vulkan backend.
model is loaded on both backends with
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --no-mmap --temp 0.1 --repeat-penalty 1.0 --top-k 20 --top-p 0.95 --min-p 0 --chat-template-kwargs '{"preserve_thinking":true,"reasoning_effort":"low"}'
Hypothesis: the issue may be context-size related — MCP tool schemas significantly
increase the prompt token count, and there may be a ROCm-specific behavior difference
under high context load compared to Vulkan.
First Bad Commit
No response
Relevant log output
I don't have any, let me know if lemonade server logs are enough, I can provide them.