Skip to content

ggml: refactor selective expert copying to user code - #29943

Merged
ggerganov merged 6 commits into
masterfrom
aman/sched-copy-callback
Oct 6, 2026
Merged

ggerganov merged 6 commits into
masterfrom
aman/sched-copy-callback

Conversation

@am17an

@am17an am17an commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Overview

Refactor the selective expert copying added in #15346 to be called from user code. No functional change is expected from this change. The idea is that this will help in #29887 to be complety handed in user-code. After this is merged, the only change there should be in user code.

Additional information

Requirements

@am17an
am17an requested a review from ggerganov as a code owner October 4, 2026 10:20
@github-actions github-actions Bot added the ggml changes relating to the ggml tensor library for machine learning label Oct 4, 2026
@ggerganov ggerganov self-assigned this Oct 4, 2026
Comment thread ggml/include/ggml-backend.h Outdated
Comment thread ggml/include/ggml-backend.h Outdated
Comment thread src/llama-context.cpp
const int64_t n_expert = src->ne[2];
const size_t expert_size = src->nb[2];

if (ids != st.ids || (int64_t) st.used.size() != n_expert) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In which situations, the ids == st.ids?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see that this case is also handled on master, but it is not clear to me why.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's because we transfer gate, up and down and they can all re-use the same ids tensors

Comment thread src/llama-context.cpp Outdated
@github-actions github-actions Bot added the testing Everything test related label Oct 4, 2026
@am17an
am17an force-pushed the aman/sched-copy-callback branch from 7441797 to 72cb4da Compare October 4, 2026 15:46
Comment thread ggml/include/ggml-backend.h
@ggerganov
ggerganov force-pushed the aman/sched-copy-callback branch from 528e0a3 to ca2934f Compare October 6, 2026 06:50
Comment thread ggml/include/ggml-backend.h Outdated
Comment thread ggml/include/ggml-backend.h Outdated
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
@ggerganov
ggerganov merged commit 6753a03 into master Oct 6, 2026
26 of 27 checks passed
@ggerganov
ggerganov deleted the aman/sched-copy-callback branch October 6, 2026 07:59
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
* ggml: refactor selective expert copying to user code

* tests: enroll two models into selective expert copy test

* tests: use deepseek2 as test model

* improve comment in ggml-backend.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* cont: fix whitespace

* cont : better comments

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
(cherry picked from commit 6753a03)
sodre90 added a commit to sodre90/llama.cpp that referenced this pull request Oct 8, 2026
sodre90 added a commit to sodre90/llama.cpp that referenced this pull request Oct 8, 2026
…ts ahead and with O_DIRECT again

The ggml-org#29943 port moved the selective expert upload into sched_copy_experts but left the lookahead source and the staged direct reads behind in the scheduler, where nothing called them. Prefill then copied every non-pool expert from the cold mapping.
kmbandy added a commit to kmbandy/llama.cpp that referenced this pull request Oct 8, 2026
Brings LiquidAI d1-3B / d1-omni-600M decision models (ggml-org#30110, ggml-org#30114),
CLEF vision, pplx-decider, LFM2.5 encoder, split router child stdout/stderr
(ggml-org#29895), MoE expert-copy callback (ggml-org#29943) and general upstream fixes.

27 conflicted files (69 hunks) resolved keeping both sides; fork-line audit
on every conflicted file and on the 57 auto-merged co-touched files: no fork
lines lost. server-node.cpp ported to the split subproc API. Server keeps the
fork's llama_batch decode path and routes decision/joint-head batches through
llama_process (common_batch). Fork's non-fatal spec/checkpoint restore kept.
Notes: ~/.cache/llama-sync-20261008/ (reports, review, rulings).
Backup: backup/pre-upstream-sync-2026-10-08 @ 49e0f85. Not yet compiled.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
mclrcha added a commit to mclrcha/llama.cpp that referenced this pull request Oct 10, 2026
…-10-10)

Conflicts:
- ggml-backend.cpp: upstream moved input copies to ggml_backend_sched_copy_input with host weights last (ggml-org#29943);
  the GGML_SCHED_BATCH_INPUTS fast path now runs in the first loop before falling back to it.
- ggml-cuda/fattn-common.cuh: launch_fattn takes upstream async_kv_preload, then our parallel_blocks_fixed
  (GQA-6 tile launcher updated).
- ggml-cuda/fattn-mma-f16.cuh: upstream swizzle refactor (ggml-org#29612); its helpers take ncols2 like the rest of
  the fork's config getters.
- ggml-cuda/gated_delta_net.cu: keep the fork's kernel and launch geometry (its columns_per_block template
  argument is not upstream's cols_per_warp, ggml-org#30087).
- ggml-cuda/mmq.cu: upstream per-tile src1 padding (ggml-org#29953); ggml_cuda_mul_mat_q_src1_nbytes and the MMQ input
  cache use the worst-case padding so producers and projections with other tiles can share the buffer.
- ggml-cuda/mmq-vec-dot.cuh: fork vec_dot kernels pass GGML_PREC_Q8 to the config helpers (ggml-org#30168).
- ggml-cuda/mmvf.cu/.cuh: keep both includes and the HIP declarations, upstream warp_size argument.
- ggml-cuda/rope.cu: upstream grid-stride rms_norm+mul+rope (ggml-org#28175) with the fork's M-RoPE theta (nchannels
  instead of gridDim.y) and rope_store_cast.
- ggml-cuda/top-k.cu: upstream shape-based selection (ggml-org#28713), RDNA4 tournament kept for few long rows.
- src/llama-batch.*: mixed token/embd batches (ggml-org#29622) copy embd rows per entry; other batches keep pointing at
  the ext rows.
- src/llama-context.cpp: expert copy callback also on the second MTP scheduler.
- src/llama-graph.cpp: build_rs keeps the GET_ROWS of the ubatch states (GDN fusions, deferred state) and gathers
  the extra states from the tail like upstream's custom getter path (ggml-org#29856).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants