Skip to content

sampling : delegate input allocation to the scheduler - #19266

Merged
ggerganov merged 2 commits into
masterfrom
gg/backend-sampling-fix-inp-allocation
Feb 3, 2026
Merged

ggerganov merged 2 commits into
masterfrom
gg/backend-sampling-fix-inp-allocation

Conversation

@ggerganov

@ggerganov ggerganov commented Feb 2, 2026 •

Copy link
Copy Markdown
Member

fix #18622
alt #18636

  • Merge the sampler inputs into the main graph. This way the backend scheduler is responsible for allocating the memory which makes backend sampling compatible with pipeline parallelism
  • Utilize ggml_build_forward_select() in llm_graph_context::build_sampling() to avoid computing the samplers when not needed

@ggerganov
ggerganov force-pushed the gg/backend-sampling-fix-inp-allocation branch from c4d5b0f to 7f58cca Compare February 3, 2026 11:01
@ggerganov ggerganov mentioned this pull request Feb 3, 2026
1 task
@ggerganov
ggerganov marked this pull request as ready for review February 3, 2026 14:12
@ggerganov
ggerganov requested a review from CISC as a code owner February 3, 2026 14:12
@ggerganov
ggerganov requested a review from danbev February 3, 2026 14:13
@ggerganov
ggerganov merged commit faa1bc2 into master Feb 3, 2026
76 of 78 checks passed
@ggerganov
ggerganov deleted the gg/backend-sampling-fix-inp-allocation branch February 3, 2026 20:16
liparetejas pushed a commit to liparetejas/llama.cpp that referenced this pull request Feb 23, 2026
* sampling : delegate input allocation to the scheduler

* graph : compute backend samplers only if needed
NihilDigit pushed a commit to NihilDigit/llama.cpp that referenced this pull request Apr 12, 2026
* sampling : delegate input allocation to the scheduler

* graph : compute backend samplers only if needed
Seunghhon pushed a commit to Seunghhon/llama.cpp that referenced this pull request Apr 26, 2026
* sampling : delegate input allocation to the scheduler

* graph : compute backend samplers only if needed
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
* sampling : delegate input allocation to the scheduler

* graph : compute backend samplers only if needed
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
* sampling : delegate input allocation to the scheduler

* graph : compute backend samplers only if needed
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
* sampling : delegate input allocation to the scheduler

* graph : compute backend samplers only if needed
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
* sampling : delegate input allocation to the scheduler

* graph : compute backend samplers only if needed
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* sampling : delegate input allocation to the scheduler

* graph : compute backend samplers only if needed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval bug: Segmentation fault with -bs with multiple GPUs

2 participants