Skip to content

CUDA: pass src1 precision to host MMQ config helpers - #30168

Merged
ynankani merged 2 commits into
ggml-org:masterfrom
ynankani:ynankani/mmq-host-prec-src1
Oct 9, 2026
Merged

ynankani merged 2 commits into
ggml-org:masterfrom
ynankani:ynankani/mmq-host-prec-src1

Conversation

@ynankani

@ynankani ynankani commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Thread prec_src1 through host-side MMQ helpers to keep host and device configuration selection consistent .

Additional information

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes, for mechanical changes at multiple places

Signed-off-by: ynankani <ynankani@nvidia.com>
@ynankani
ynankani requested a review from a team as a code owner October 8, 2026 17:38
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Oct 8, 2026

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Preferably also remove the default of GGML_PREC_Q8 from all functions to avoid accidental misuse.

Signed-off-by: ynankani <ynankani@nvidia.com>
@ynankani

ynankani commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Preferably also remove the default of GGML_PREC_Q8 from all functions to avoid accidental misuse.

Yes, it makes sense to supply it src1 prec explicitly

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you.

@JohannesGaessler JohannesGaessler added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Oct 9, 2026
@ynankani
ynankani merged commit 8a1a9b5 into ggml-org:master Oct 9, 2026
16 of 18 checks passed
mclrcha added a commit to mclrcha/llama.cpp that referenced this pull request Oct 10, 2026
…-10-10)

Conflicts:
- ggml-backend.cpp: upstream moved input copies to ggml_backend_sched_copy_input with host weights last (ggml-org#29943);
  the GGML_SCHED_BATCH_INPUTS fast path now runs in the first loop before falling back to it.
- ggml-cuda/fattn-common.cuh: launch_fattn takes upstream async_kv_preload, then our parallel_blocks_fixed
  (GQA-6 tile launcher updated).
- ggml-cuda/fattn-mma-f16.cuh: upstream swizzle refactor (ggml-org#29612); its helpers take ncols2 like the rest of
  the fork's config getters.
- ggml-cuda/gated_delta_net.cu: keep the fork's kernel and launch geometry (its columns_per_block template
  argument is not upstream's cols_per_warp, ggml-org#30087).
- ggml-cuda/mmq.cu: upstream per-tile src1 padding (ggml-org#29953); ggml_cuda_mul_mat_q_src1_nbytes and the MMQ input
  cache use the worst-case padding so producers and projections with other tiles can share the buffer.
- ggml-cuda/mmq-vec-dot.cuh: fork vec_dot kernels pass GGML_PREC_Q8 to the config helpers (ggml-org#30168).
- ggml-cuda/mmvf.cu/.cuh: keep both includes and the HIP declarations, upstream warp_size argument.
- ggml-cuda/rope.cu: upstream grid-stride rms_norm+mul+rope (ggml-org#28175) with the fork's M-RoPE theta (nchannels
  instead of gridDim.y) and rope_store_cast.
- ggml-cuda/top-k.cu: upstream shape-based selection (ggml-org#28713), RDNA4 tournament kept for few long rows.
- src/llama-batch.*: mixed token/embd batches (ggml-org#29622) copy embd rows per entry; other batches keep pointing at
  the ext rows.
- src/llama-context.cpp: expert copy callback also on the second MTP scheduler.
- src/llama-graph.cpp: build_rs keeps the GET_ROWS of the ubatch states (GDN fusions, deferred state) and gathers
  the extra states from the tail like upstream's custom getter path (ggml-org#29856).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants