Repository navigation
Conversation
std::string mesh up vocab.
AAbushady
pushed a commit
to AAbushady/llama.cpp
that referenced
this pull request
Jan 27, 2024
* Initial conversion to customtkinter. * Initial conversion to customtkinter. * Additions to UI, still non-functional * UI now functional, untested * UI now functional, untested * Added saving configs * Saving and loading now functional * Fixed sliders not loading * Cleaned up duplicate arrays * Cleaned up duplicate arrays * Fixed loading bugs * wip fixing all the broken parameters. PLEASE test before you commit * further cleaning * bugfix completed for gui. now evaluating save and load * cleanup prepare to merge --------- Co-authored-by: Concedo <39025047+LostRuins@users.noreply.github.com>
phuongncn
pushed a commit
to phuongncn/llama.cpp-gx10-dgx-sparks-deepseekv4
that referenced
this pull request
Apr 28, 2026
…gml-org#284) * llama-bench: enable having different number of threads for tg and pp * Add -tgb to usage --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
This was referenced Oct 2, 2026
bong-water-water-bong
added a commit
to 1bit-MONSTER/llama.cpp
that referenced
this pull request
Oct 5, 2026
ggml-hrx: the f32_f32 WMMA cores accumulate in f32 (engine ggml-org#284)
Copilot AI
pushed a commit
to imaami/llama.cpp
that referenced
this pull request
Oct 6, 2026
…gml-org#284) * dflash: add beam selection and reuse Metal verify data * metal: few-row tensor tile for PQ2_0 at speculative verify widths (ggml-org#286) Adds kernel_mul_mm_pq2_0_f32_fewrow, a register-resident matmul2d tile of 16 src1 rows by 32 weight rows per simdgroup. PQ2_0 weights are decoded straight into the right-hand cooperative tensor, so the K loop needs no threadgroup staging and no barriers. On Metal 4 devices the generic tensor mul_mm tile is 64 x 128, which computes far more columns than a 4..16 row verify batch needs, and mul_mv_ext is ALU-bound at 2..8 columns. The route is taken only when the device has the tensor API, src0 is PQ2_0, src1 is contiguous F32, there is no batching in dims 2/3, and ne11 is in [GGML_METAL_PQ2_0_FEWROW_MIN, _MAX] (default 3..32). Below 3 columns the existing mat-vec paths stay faster. GGML_METAL_PQ2_0_FEWROW=0 turns the route off, and GGML_METAL_PQ2_0_FEWROW_CFG selects the column-block by K-split layout. test-backend-ops perf on M5 Pro, 17408 x 5120, best of 2 (us/run): n=3 187.0 -> 151.3 (3-column mul_mv from the parent commit) n=8 484.5 -> 157.4 (mul_mv_ext) n=16 855.9 -> 172.8 (tensor mul_mm) n=32 864.0 -> 328.3 n=1..2 are slower on the new tile (150 vs 71 and 130 us), hence min 3. End to end on M5 Pro (27B PQ2_0 target, DFlash2 drafter, greedy, depth 7, 8 prompts x 3 reps): with this route off, speculation runs at 0.73x of plain decoding; with it on, 1.57x (BF16 drafter) to 1.64x (Q8_0 drafter). Output matches plain decoding except one near-tie (top-2 logprob gap 0.00047) where the half-precision activations pick the other token. The cooperative tensor element layout was probed on M5 with get_multidimensional_index before writing the decode. test-backend-ops gains PQ2_0 MUL_MAT cases for every n from 1 to 32 plus 64 and 512, and perf cases at the 27B projection shapes, plus flash attention perf cases at the DFlash2 drafter and 27B verify shapes. The few-row approach is adapted from the qmm_m16_block core in the Yukon MLX.fast promoted submissions (polymorf and later contributors), used under Apache-2.0. The PQ2_0 block layout differs, so the decode and dispatch are new code. (cherry picked from commit 3285e75) Assisted-by: GitHub Copilot
Copilot AI
pushed a commit
to imaami/llama.cpp
that referenced
this pull request
Oct 6, 2026
…gml-org#284) * dflash: add beam selection and reuse Metal verify data * metal: few-row tensor tile for PQ2_0 at speculative verify widths (ggml-org#286) Adds kernel_mul_mm_pq2_0_f32_fewrow, a register-resident matmul2d tile of 16 src1 rows by 32 weight rows per simdgroup. PQ2_0 weights are decoded straight into the right-hand cooperative tensor, so the K loop needs no threadgroup staging and no barriers. On Metal 4 devices the generic tensor mul_mm tile is 64 x 128, which computes far more columns than a 4..16 row verify batch needs, and mul_mv_ext is ALU-bound at 2..8 columns. The route is taken only when the device has the tensor API, src0 is PQ2_0, src1 is contiguous F32, there is no batching in dims 2/3, and ne11 is in [GGML_METAL_PQ2_0_FEWROW_MIN, _MAX] (default 3..32). Below 3 columns the existing mat-vec paths stay faster. GGML_METAL_PQ2_0_FEWROW=0 turns the route off, and GGML_METAL_PQ2_0_FEWROW_CFG selects the column-block by K-split layout. test-backend-ops perf on M5 Pro, 17408 x 5120, best of 2 (us/run): n=3 187.0 -> 151.3 (3-column mul_mv from the parent commit) n=8 484.5 -> 157.4 (mul_mv_ext) n=16 855.9 -> 172.8 (tensor mul_mm) n=32 864.0 -> 328.3 n=1..2 are slower on the new tile (150 vs 71 and 130 us), hence min 3. End to end on M5 Pro (27B PQ2_0 target, DFlash2 drafter, greedy, depth 7, 8 prompts x 3 reps): with this route off, speculation runs at 0.73x of plain decoding; with it on, 1.57x (BF16 drafter) to 1.64x (Q8_0 drafter). Output matches plain decoding except one near-tie (top-2 logprob gap 0.00047) where the half-precision activations pick the other token. The cooperative tensor element layout was probed on M5 with get_multidimensional_index before writing the decode. test-backend-ops gains PQ2_0 MUL_MAT cases for every n from 1 to 32 plus 64 and 512, and perf cases at the 27B projection shapes, plus flash attention perf cases at the DFlash2 drafter and 27B verify shapes. The few-row approach is adapted from the qmm_m16_block core in the Yukon MLX.fast promoted submissions (polymorf and later contributors), used under Apache-2.0. The PQ2_0 block layout differs, so the decode and dispatch are new code. (cherry picked from commit 3285e75) Assisted-by: GitHub Copilot
imaami
pushed a commit
to imaami/llama.cpp
that referenced
this pull request
Oct 6, 2026
…gml-org#284) * dflash: add beam selection and reuse Metal verify data * metal: few-row tensor tile for PQ2_0 at speculative verify widths (ggml-org#286) Adds kernel_mul_mm_pq2_0_f32_fewrow, a register-resident matmul2d tile of 16 src1 rows by 32 weight rows per simdgroup. PQ2_0 weights are decoded straight into the right-hand cooperative tensor, so the K loop needs no threadgroup staging and no barriers. On Metal 4 devices the generic tensor mul_mm tile is 64 x 128, which computes far more columns than a 4..16 row verify batch needs, and mul_mv_ext is ALU-bound at 2..8 columns. The route is taken only when the device has the tensor API, src0 is PQ2_0, src1 is contiguous F32, there is no batching in dims 2/3, and ne11 is in [GGML_METAL_PQ2_0_FEWROW_MIN, _MAX] (default 3..32). Below 3 columns the existing mat-vec paths stay faster. GGML_METAL_PQ2_0_FEWROW=0 turns the route off, and GGML_METAL_PQ2_0_FEWROW_CFG selects the column-block by K-split layout. test-backend-ops perf on M5 Pro, 17408 x 5120, best of 2 (us/run): n=3 187.0 -> 151.3 (3-column mul_mv from the parent commit) n=8 484.5 -> 157.4 (mul_mv_ext) n=16 855.9 -> 172.8 (tensor mul_mm) n=32 864.0 -> 328.3 n=1..2 are slower on the new tile (150 vs 71 and 130 us), hence min 3. End to end on M5 Pro (27B PQ2_0 target, DFlash2 drafter, greedy, depth 7, 8 prompts x 3 reps): with this route off, speculation runs at 0.73x of plain decoding; with it on, 1.57x (BF16 drafter) to 1.64x (Q8_0 drafter). Output matches plain decoding except one near-tie (top-2 logprob gap 0.00047) where the half-precision activations pick the other token. The cooperative tensor element layout was probed on M5 with get_multidimensional_index before writing the decode. test-backend-ops gains PQ2_0 MUL_MAT cases for every n from 1 to 32 plus 64 and 512, and perf cases at the 27B projection shapes, plus flash attention perf cases at the DFlash2 drafter and 27B verify shapes. The few-row approach is adapted from the qmm_m16_block core in the Yukon MLX.fast promoted submissions (polymorf and later contributors), used under Apache-2.0. The PQ2_0 block layout differs, so the decode and dispatch are new code. (cherry picked from commit 3285e75) Assisted-by: GitHub Copilot
imaami
pushed a commit
to imaami/llama.cpp
that referenced
this pull request
Oct 6, 2026
…gml-org#284) * dflash: add beam selection and reuse Metal verify data * metal: few-row tensor tile for PQ2_0 at speculative verify widths (ggml-org#286) Adds kernel_mul_mm_pq2_0_f32_fewrow, a register-resident matmul2d tile of 16 src1 rows by 32 weight rows per simdgroup. PQ2_0 weights are decoded straight into the right-hand cooperative tensor, so the K loop needs no threadgroup staging and no barriers. On Metal 4 devices the generic tensor mul_mm tile is 64 x 128, which computes far more columns than a 4..16 row verify batch needs, and mul_mv_ext is ALU-bound at 2..8 columns. The route is taken only when the device has the tensor API, src0 is PQ2_0, src1 is contiguous F32, there is no batching in dims 2/3, and ne11 is in [GGML_METAL_PQ2_0_FEWROW_MIN, _MAX] (default 3..32). Below 3 columns the existing mat-vec paths stay faster. GGML_METAL_PQ2_0_FEWROW=0 turns the route off, and GGML_METAL_PQ2_0_FEWROW_CFG selects the column-block by K-split layout. test-backend-ops perf on M5 Pro, 17408 x 5120, best of 2 (us/run): n=3 187.0 -> 151.3 (3-column mul_mv from the parent commit) n=8 484.5 -> 157.4 (mul_mv_ext) n=16 855.9 -> 172.8 (tensor mul_mm) n=32 864.0 -> 328.3 n=1..2 are slower on the new tile (150 vs 71 and 130 us), hence min 3. End to end on M5 Pro (27B PQ2_0 target, DFlash2 drafter, greedy, depth 7, 8 prompts x 3 reps): with this route off, speculation runs at 0.73x of plain decoding; with it on, 1.57x (BF16 drafter) to 1.64x (Q8_0 drafter). Output matches plain decoding except one near-tie (top-2 logprob gap 0.00047) where the half-precision activations pick the other token. The cooperative tensor element layout was probed on M5 with get_multidimensional_index before writing the decode. test-backend-ops gains PQ2_0 MUL_MAT cases for every n from 1 to 32 plus 64 and 512, and perf cases at the 27B projection shapes, plus flash attention perf cases at the DFlash2 drafter and 27B verify shapes. The few-row approach is adapted from the qmm_m16_block core in the Yukon MLX.fast promoted submissions (polymorf and later contributors), used under Apache-2.0. The PQ2_0 block layout differs, so the decode and dispatch are new code. (cherry picked from commit 3285e75) Assisted-by: GitHub Copilot
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

bugfix: std::string mesh up vocab.
OS: CentOS 7
compiler: gcc (GCC) 11.2.1 20220127 (Red Hat 11.2.1-9)