Skip to content

ggml: add SWIGLU_CLAMP - #27930

Merged
am17an merged 2 commits into
ggml-org:masterfrom
am17an:swiglu_clamp
Aug 30, 2026
Merged

am17an merged 2 commits into
ggml-org:masterfrom
am17an:swiglu_clamp

Conversation

@am17an

@am17an am17an commented Aug 29, 2026

Copy link
Copy Markdown
Contributor

Overview

DSV4, GLM use a new variant of GLU which clamps the SWIGLU. Currently in the graph it is not represented well for downstream fusion and creates 3 separate OPs. Helps 2% in TG and 1-2% on PP, tested on 4x4090s. It should help in spec-dec as well after we merge #27621

Additional information

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, I wrote some of the CUDA and llama-graph changes. Rest of the code is generated by AI and made to pass test-backend-ops wherever possible

@am17an
am17an requested review from a team, CISC, ggerganov and marty1885 as code owners August 29, 2026 06:28
@github-actions github-actions Bot added testing Everything test related Vulkan Issues specific to the Vulkan backend ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language Apple Metal https://en.wikipedia.org/wiki/Metal_(API) Ascend NPU issues specific to Ascend NPUs OpenCL Issues specific to the OpenCL backend Hexagon CUDA Related to the CUDA backend OpenVINO WebGPU labels Aug 29, 2026
@CISC

CISC commented Aug 29, 2026

Copy link
Copy Markdown
Member

Can't we just add this as a universal glu op param instead of making a new op?

@am17an

am17an commented Aug 29, 2026

Copy link
Copy Markdown
Contributor Author

It's not an op, it's added to glu enum right?

@CISC

CISC commented Aug 29, 2026

Copy link
Copy Markdown
Member

It's not an op, it's added to glu enum right?

Sure, still. :)

@am17an

am17an commented Aug 29, 2026

Copy link
Copy Markdown
Contributor Author

I'm okay either way, it would not lead to less changes and it would complicate at least the fusion in the CUDA backend. We already have a multiple glu variants and this is quite a valid one which is used is more than 1 arch.

@ggerganov

Copy link
Copy Markdown
Member

universal glu op param

Could you clarify? The swiglu+clamp also needs a limit parameter.

@CISC

CISC commented Aug 29, 2026

Copy link
Copy Markdown
Member

universal glu op param

Could you clarify? The swiglu+clamp also needs a limit parameter.

I meant that all glu ops could check op param 3 for limit.

@jeffbolznv jeffbolznv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ggml-vulkan changes look good.

@max-krasnyansky max-krasnyansky left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hexagon changes look good.
opencl looks good too, @lhez please double-check

@marty1885 marty1885 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval for the ET portion

@0cc4m

0cc4m commented Aug 30, 2026

Copy link
Copy Markdown
Contributor

Vulkan code is fine.

@am17an

am17an commented Aug 30, 2026

Copy link
Copy Markdown
Contributor Author

@CISC can I merge?

@CISC

CISC commented Aug 30, 2026

Copy link
Copy Markdown
Member

@CISC can I merge?

Sure, it was merely a suggestion to avoid an extra enum and possibly support future variations.

@am17an
am17an merged commit 0190529 into ggml-org:master Aug 30, 2026
52 of 55 checks passed
@am17an
am17an deleted the swiglu_clamp branch August 30, 2026 15:00
@Crandel

Crandel commented Aug 31, 2026

Copy link
Copy Markdown

Failed to build llama.cpp after this PR was merged

FAILED: [code=1] src/CMakeFiles/llama.dir/llama-graph.cpp.o
/usr/bin/c++ -DLLAMA_BUILD -DLLAMA_COMMIT=\"daef7b6874\" -DLLAMA_SHARED -DLLAMA_VERSION=\"0.3.0-dev\" -Dllama_EXPORTS -I/data/work/projects/aur/llama-cpp/src/llama.cpp/src/. -I/data/work/projects/aur/llama-cpp/src/llama.cpp/src/../include -march=x86-64 -mtune=generic -O2 -pipe -fno-plt -fexceptions         -Wp,-D_FORTIFY_SOURCE=3 -Wformat -Werror=format-security         -fstack-clash-protection -fcf-protection         -fno-omit-frame-pointer -mno-omit-leaf-frame-pointer -Wp,-D_GLIBCXX_ASSERTIONS -flto=auto -fPIC -Wmissing-declarations -Wmissing-noreturn -Wall -Wextra -Wpedantic -Wcast-qual -Wno-unused-function -Wno-array-bounds -Wextra-semi -MD -MT src/CMakeFiles/llama.dir/llama-graph.cpp.o -MF src/CMakeFiles/llama.dir/llama-graph.cpp.o.d -o src/CMakeFiles/llama.dir/llama-graph.cpp.o -c /data/work/projects/aur/llama-cpp/src/llama.cpp/src/llama-graph.cpp
/data/work/projects/aur/llama-cpp/src/llama.cpp/src/llama-graph.cpp: In member function ‘ggml_tensor* llm_graph_context::build_ffn(ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, llm_ffn_op_type, llm_ffn_gate_type, int) const’:
/data/work/projects/aur/llama-cpp/src/llama.cpp/src/llama-graph.cpp:1780:35: error: ‘ggml_swiglu_clamp’ was not declared in this scope; did you mean ‘ggml_swiglu_oai’?
 1780 |                             cur = ggml_swiglu_clamp(ctx0, cur, tmp, limit);
      |                                   ^~~~~~~~~~~~~~~~~
      |                                   ggml_swiglu_oai
/data/work/projects/aur/llama-cpp/src/llama.cpp/src/llama-graph.cpp: In member function ‘ggml_tensor* llm_graph_context::build_moe_ffn(ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, int64_t, int64_t, llm_ffn_op_type, bool, float, llama_expert_gating_func_type, int, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*, ggml_tensor*) const’:
/data/work/projects/aur/llama-cpp/src/llama.cpp/src/llama-graph.cpp:2174:35: error: ‘ggml_swiglu_clamp’ was not declared in this scope; did you mean ‘ggml_swiglu_oai’?
 2174 |                             cur = ggml_swiglu_clamp(ctx0, cur, up, limit);
      |                                   ^~~~~~~~~~~~~~~~~
      |                                   ggml_swiglu_oai
[39/400] Building CXX object src/CMakeFiles/llama.dir/unicode.cpp.o

Arch build script

  cmake -S llama.cpp -B build -G Ninja \
      -DCMAKE_BUILD_TYPE=None \
      -DCMAKE_INSTALL_PREFIX=/usr \
      -DLLAMA_BUILD_APP=OFF \
      -DLLAMA_BUILD_EXAMPLES=ON \
      -DLLAMA_BUILD_SERVER=ON \
      -DLLAMA_BUILD_TESTS=OFF \
      -DLLAMA_BUILD_TOOLS=ON \
      -DLLAMA_BUILD_UI=ON \
      -DLLAMA_USE_SYSTEM_GGML=ON
  cmake --build build

@ggerganov

ggerganov commented Aug 31, 2026 •

Copy link
Copy Markdown
Member

-DLLAMA_USE_SYSTEM_GGML=ON

llama.cpp has a local copy of ggml - you should not use a system-wide ggml with development versions of llama.cpp. You can use only one with the stable versions (e.g. v0.3.0). More info: ggml-org/ggml#1579

@Crandel

Crandel commented Aug 31, 2026

Copy link
Copy Markdown

llama.cpp has a local copy of ggml - you should not use a system-wide ggml with development versions of llama.cpp. You can use only one with the stable versions (e.g. v0.3.0). More info: ggml-org/ggml#1579

Thank you for help, I didn't know about ggml separation before. Then I will wait for next release to test Qwen-3.8-Flash, as I was trying to update systemwide llama.cpp package.

@fairydreaming

Copy link
Copy Markdown
Contributor

I noticed that OpenVINO backend CI fails some new tests https://github.com/ggml-org/llama.cpp/actions/runs/33609523462/job/100181219768

2026-09-02T10:52:59.9409704Z [SWIGLU_CLAMP] ERR = 0.000000307 > 0.000000100   SWIGLU_CLAMP(type=f16,ne_a=[128,2,2,2],v=0,limit=2.000000): �[1;31mFAIL�[0m
2026-09-02T10:52:59.9410165Z   SWIGLU_CLAMP(type=f16,ne_a=[128,2,2,2],v=0,limit=10.000000): �[1;32mOK�[0m
2026-09-02T10:52:59.9410636Z [SWIGLU_CLAMP] ERR = 0.000000307 > 0.000000100   SWIGLU_CLAMP(type=f16,ne_a=[128,2,2,2],v=1,limit=2.000000): �[1;31mFAIL�[0m
2026-09-02T10:52:59.9411084Z   SWIGLU_CLAMP(type=f16,ne_a=[128,2,2,2],v=1,limit=10.000000): �[1;32mOK�[0m
2026-09-02T10:52:59.9411448Z   SWIGLU_CLAMP(type=f32,ne_a=[128,2,2,2],v=0,limit=2.000000): �[1;32mOK�[0m
2026-09-02T10:52:59.9411798Z   SWIGLU_CLAMP(type=f32,ne_a=[128,2,2,2],v=0,limit=10.000000): �[1;32mOK�[0m
2026-09-02T10:52:59.9412144Z   SWIGLU_CLAMP(type=f32,ne_a=[128,2,2,2],v=1,limit=2.000000): �[1;32mOK�[0m
2026-09-02T10:52:59.9412488Z   SWIGLU_CLAMP(type=f32,ne_a=[128,2,2,2],v=1,limit=10.000000): �[1;32mOK�[0m

err looks low, perhaps this test needs increased max_nmse_err()?

@ggerganov

Copy link
Copy Markdown
Member

AFAIK, the OpenVINO team will be addressing these soon.

@BrewTestBot BrewTestBot mentioned this pull request Sep 4, 2026
1 task done
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
* ggml: add SWIGLU_CLAMP

* add vulkan shader
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* ggml: add SWIGLU_CLAMP

* add vulkan shader
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* ggml: add SWIGLU_CLAMP

* add vulkan shader
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* ggml: add SWIGLU_CLAMP

* add vulkan shader
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) Ascend NPU issues specific to Ascend NPUs CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning Hexagon OpenCL Issues specific to the OpenCL backend OpenVINO SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language testing Everything test related Vulkan Issues specific to the Vulkan backend WebGPU

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants