Skip to content

sycl: stage bulk uploads (model loading) through a pinned ring buffer - #29608

Merged
Titaniumtown merged 1 commit into
ggml-org:masterfrom
cwriter:pr/09-sycl-pinned-upload-ring
Oct 8, 2026
Merged

Titaniumtown merged 1 commit into
ggml-org:masterfrom
cwriter:pr/09-sycl-pinned-upload-ring

Conversation

@cwriter

@cwriter cwriter commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Overview

SYCL's upload of mmap'ed tensors currently synchronizes on a buffer allocated and freed per-transfer. This is extremely slow when loading big models.

This patch introduces a staging ringbuffer instead, which allows transfers to be prepared and sent while a previous transfer completes. The reuse of the buffer (rather than freshly allocating) further reduces the overhead.

Additional information

The effects only show for load-mode mmap on large models. For me, it's unsloth/qwen3.8-flash-next:IQ4_XXS on 3x Arc Pro B60 (each on x8 PCIe 3.0) on an x399 Threadripper platform. The env var GGML_SYCL_UPLOAD_STAGING_SLOTS can be used to disable this behavior (by setting =0). Defaults to 4 staging buffers as a trade off between required memory and saturating the copy engines.

Slots load time
0 (master) 96.8s
4 (this PR) 47.8s

This has no effect on PP and TG. It's purely a speedup for loading the model; it therefore reduces start to "slots idle" in llama-server.

Currently a draft due to the open PR limit.

Requirements

  • I have read and agree with the contributing guidelines - YES
  • AI usage disclosure: YES, used claude opus 5.5 to implement. I reviewed and created the PR manually.

@github-actions github-actions Bot added documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language labels Sep 28, 2026

@arthw arthw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's good job!

It can reduce the load time.
Here is the test result on Arc770:

CMD: time ./examples/sycl/test.sh -m ../models/gpt-oss-20b-mxfp4.gguf

Metric Base (Baseline) PR (Pull Request) Absolute Change Relative Change (Reduction)
Real Time (Wall clock) 24.314s 17.429s -6.885s -28.32%
User Time (CPU in user space) 11.201s 10.651s -0.550s -4.91%
Sys Time (CPU in kernel space) 13.387s 7.085s -6.302s -47.08%

No impact of PP and TG.
Here is my test result on Arc770:

GGUF Test fa Base t/s Primary t/s Increase Rate (Primary vs Base)
Qwen3.5-4B-Q4_K_M.gguf pp512 0 1331.13 1332.02 0.07%
Qwen3.5-4B-Q4_K_M.gguf pp512 1 1350.39 1350.93 0.04%
Qwen3.5-4B-Q4_K_M.gguf tg128 0 32.18 32.02 -0.50%
Qwen3.5-4B-Q4_K_M.gguf tg128 1 32.82 32.56 -0.79%
GGUF Test fa Base t/s Primary t/s Increase Rate (Primary vs Base)
Qwen3-30B-A3B-UD-IQ3_XXS.gguf pp512 0 294.80 294.67 -0.04%
Qwen3-30B-A3B-UD-IQ3_XXS.gguf pp512 1 299.36 300.26 0.30%
Qwen3-30B-A3B-UD-IQ3_XXS.gguf tg128 0 16.69 16.45 -1.44%
Qwen3-30B-A3B-UD-IQ3_XXS.gguf tg128 1 19.76 19.91 0.76%
GGUF Test fa Base t/s Primary t/s Increase Rate (Primary vs Base)
gpt-oss-20b-mxfp4.gguf pp512 0 786.72 787.00 0.04%
gpt-oss-20b-mxfp4.gguf pp512 1 754.40 752.61 -0.24%
gpt-oss-20b-mxfp4.gguf tg128 0 20.26 20.49 1.14%
gpt-oss-20b-mxfp4.gguf tg128 1 23.85 23.86 0.04%

@arthw arthw added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Oct 7, 2026
@cwriter
cwriter marked this pull request as ready for review October 8, 2026 04:59
@cwriter
cwriter requested a review from a team as a code owner October 8, 2026 04:59
@Titaniumtown
Titaniumtown merged commit 000bee5 into ggml-org:master Oct 8, 2026
16 of 17 checks passed
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
…ggml-org#29608)

Co-authored-by: cwriter <cwriter@localhost>
(cherry picked from commit 000bee5)
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 9, 2026
…and gate fusion

Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU,
ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits
before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the
upstream side back where both belong:

- common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary
  predicate, test_top_k_inf)
- mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column
  dispatch is upstream's (weights shared and rows paired on Xe2)
- ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it
- the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and
  keeps this branch's rule that pairs both reordered run unfused
- MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it

Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs
each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 /
200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card
and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211).

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 10, 2026
…and gate fusion

Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU,
ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits
before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the
upstream side back where both belong:

- common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary
  predicate, test_top_k_inf)
- mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column
  dispatch is upstream's (weights shared and rows paired on Xe2)
- ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it
- the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and
  keeps this branch's rule that pairs both reordered run unfused
- MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it

Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs
each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 /
200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card
and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211).

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 10, 2026
…and gate fusion

Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU,
ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits
before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the
upstream side back where both belong:

- common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary
  predicate, test_top_k_inf)
- mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column
  dispatch is upstream's (weights shared and rows paired on Xe2)
- ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it
- the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and
  keeps this branch's rule that pairs both reordered run unfused
- MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it

Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs
each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 /
200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card
and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211).

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 10, 2026
…and gate fusion

Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU,
ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits
before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the
upstream side back where both belong:

- common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary
  predicate, test_top_k_inf)
- mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column
  dispatch is upstream's (weights shared and rows paired on Xe2)
- ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it
- the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and
  keeps this branch's rule that pairs both reordered run unfused
- MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it

Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs
each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 /
200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card
and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211).

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 10, 2026
…and gate fusion

Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU,
ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits
before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the
upstream side back where both belong:

- common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary
  predicate, test_top_k_inf)
- mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column
  dispatch is upstream's (weights shared and rows paired on Xe2)
- ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it
- the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and
  keeps this branch's rule that pairs both reordered run unfused
- MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it

Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs
each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 /
200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card
and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211).

Assisted-by: Claude Opus 5.5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants