Repository navigation
sycl: stage bulk uploads (model loading) through a pinned ring buffer - #29608
Merged
Titaniumtown merged 1 commit intoOct 8, 2026
Merged
Conversation
Assisted-by: Claude Opus 5.5
arthw
approved these changes
Sep 30, 2026
arthw
left a comment
Contributor
There was a problem hiding this comment.
It's good job!
It can reduce the load time.
Here is the test result on Arc770:
CMD: time ./examples/sycl/test.sh -m ../models/gpt-oss-20b-mxfp4.gguf
| Metric | Base (Baseline) | PR (Pull Request) | Absolute Change | Relative Change (Reduction) |
|---|---|---|---|---|
| Real Time (Wall clock) | 24.314s | 17.429s | -6.885s | -28.32% |
| User Time (CPU in user space) | 11.201s | 10.651s | -0.550s | -4.91% |
| Sys Time (CPU in kernel space) | 13.387s | 7.085s | -6.302s | -47.08% |
No impact of PP and TG.
Here is my test result on Arc770:
| GGUF | Test | fa | Base t/s | Primary t/s | Increase Rate (Primary vs Base) |
|---|---|---|---|---|---|
| Qwen3.5-4B-Q4_K_M.gguf | pp512 | 0 | 1331.13 | 1332.02 | 0.07% |
| Qwen3.5-4B-Q4_K_M.gguf | pp512 | 1 | 1350.39 | 1350.93 | 0.04% |
| Qwen3.5-4B-Q4_K_M.gguf | tg128 | 0 | 32.18 | 32.02 | -0.50% |
| Qwen3.5-4B-Q4_K_M.gguf | tg128 | 1 | 32.82 | 32.56 | -0.79% |
| GGUF | Test | fa | Base t/s | Primary t/s | Increase Rate (Primary vs Base) |
|---|---|---|---|---|---|
| Qwen3-30B-A3B-UD-IQ3_XXS.gguf | pp512 | 0 | 294.80 | 294.67 | -0.04% |
| Qwen3-30B-A3B-UD-IQ3_XXS.gguf | pp512 | 1 | 299.36 | 300.26 | 0.30% |
| Qwen3-30B-A3B-UD-IQ3_XXS.gguf | tg128 | 0 | 16.69 | 16.45 | -1.44% |
| Qwen3-30B-A3B-UD-IQ3_XXS.gguf | tg128 | 1 | 19.76 | 19.91 | 0.76% |
| GGUF | Test | fa | Base t/s | Primary t/s | Increase Rate (Primary vs Base) |
|---|---|---|---|---|---|
| gpt-oss-20b-mxfp4.gguf | pp512 | 0 | 786.72 | 787.00 | 0.04% |
| gpt-oss-20b-mxfp4.gguf | pp512 | 1 | 754.40 | 752.61 | -0.24% |
| gpt-oss-20b-mxfp4.gguf | tg128 | 0 | 20.26 | 20.49 | 1.14% |
| gpt-oss-20b-mxfp4.gguf | tg128 | 1 | 23.85 | 23.86 | 0.04% |
cwriter
marked this pull request as ready for review
October 8, 2026 04:59
Titaniumtown
approved these changes
Oct 8, 2026
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 8, 2026
…ggml-org#29608) Co-authored-by: cwriter <cwriter@localhost> (cherry picked from commit 000bee5)
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 9, 2026
…and gate fusion Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU, ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the upstream side back where both belong: - common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary predicate, test_top_k_inf) - mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column dispatch is upstream's (weights shared and rows paired on Xe2) - ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it - the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and keeps this branch's rule that pairs both reordered run unfused - MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 / 200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211). Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 10, 2026
…and gate fusion Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU, ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the upstream side back where both belong: - common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary predicate, test_top_k_inf) - mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column dispatch is upstream's (weights shared and rows paired on Xe2) - ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it - the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and keeps this branch's rule that pairs both reordered run unfused - MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 / 200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211). Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 10, 2026
…and gate fusion Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU, ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the upstream side back where both belong: - common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary predicate, test_top_k_inf) - mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column dispatch is upstream's (weights shared and rows paired on Xe2) - ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it - the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and keeps this branch's rule that pairs both reordered run unfused - MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 / 200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211). Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 10, 2026
…and gate fusion Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU, ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the upstream side back where both belong: - common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary predicate, test_top_k_inf) - mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column dispatch is upstream's (weights shared and rows paired on Xe2) - ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it - the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and keeps this branch's rule that pairs both reordered run unfused - MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 / 200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211). Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 10, 2026
…and gate fusion Upstream merged its own versions of work this branch carries (ggml-org#29375 Q5_K reorder-layout MMVQ and fused GLU, ggml-org#29245 grouped MoE XMX GEMM, ggml-org#29608 pinned upload staging, ggml-org#29687 delta-net alpha gate fusion), and the commits before this one were replayed onto it with their own side taken in every conflicting hunk. This commit puts the upstream side back where both belong: - common.hpp, fusion.cpp, the op tests: both sides' additions (tile schedule of the grouped GEMM, the fused unary predicate, test_top_k_inf) - mul_mat_vec_q_reorder_ncols takes both new template flags (has_tiles, shared_weights); the Q5_K multi-column dispatch is upstream's (weights shared and rows paired on Xe2) - ggml_sycl_op_mul_mat_sycl tries upstream's fused dequant GEMM first and keeps the bf16 src0 route after it - the gate / up / GLU fusion uses upstream's reorder pairs (same-type Q4_K, Q5_K on Xe2 up to 5 columns) and keeps this branch's rule that pairs both reordered run unfused - MUL_MAT_ID keeps this branch's sorted loop; upstream's grouped GEMM is built but not called from it Qwen3.8-27B on two Arc Pro B70 against the tip before the rebase (llama-bench, depth 0, three mirrored runs each): 2 to 32 rows equal or 1-4% faster (inside today's 2-4% run-to-run spread), 48 / 64 rows 190 -> 160 / 200 -> 168 ms per batch (the fused dequant GEMM), 1024 rows equal. Perplexity unchanged on the 27B (one card and two, 1024-row and 8-row batches) and on Ornith. 17/17 op suites clean (MUL_MAT 2711, FLASH_ATTN_EXT 4211). Assisted-by: Claude Opus 5.5
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
SYCL's upload of mmap'ed tensors currently synchronizes on a buffer allocated and freed per-transfer. This is extremely slow when loading big models.
This patch introduces a staging ringbuffer instead, which allows transfers to be prepared and sent while a previous transfer completes. The reuse of the buffer (rather than freshly allocating) further reduces the overhead.
Additional information
The effects only show for load-mode mmap on large models. For me, it's unsloth/qwen3.8-flash-next:IQ4_XXS on 3x Arc Pro B60 (each on x8 PCIe 3.0) on an x399 Threadripper platform. The env var
GGML_SYCL_UPLOAD_STAGING_SLOTScan be used to disable this behavior (by setting=0). Defaults to 4 staging buffers as a trade off between required memory and saturating the copy engines.This has no effect on PP and TG. It's purely a speedup for loading the model; it therefore reduces start to "slots idle" in llama-server.
Currently a draft due to the open PR limit.
Requirements