Skip to content

bugfix: centos 7, gcc (GCC) 11.2.1 20220127 (Red Hat 11.2.1-9) - #284

Closed
OvJat wants to merge 2 commits into
ggml-org:masterfrom
OvJat:master
Closed

OvJat wants to merge 2 commits into
ggml-org:masterfrom
OvJat:master

Conversation

@OvJat

@OvJat OvJat commented Mar 19, 2023

Copy link
Copy Markdown

bugfix: std::string mesh up vocab.
OS: CentOS 7
compiler: gcc (GCC) 11.2.1 20220127 (Red Hat 11.2.1-9)

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Better fix it the way we did for whisper.cpp:

image

@gjmulder gjmulder added the bug Something isn't working label Mar 20, 2023
@OvJat OvJat closed this Mar 21, 2023
AAbushady pushed a commit to AAbushady/llama.cpp that referenced this pull request Jan 27, 2024
* Initial conversion to customtkinter.

* Initial conversion to customtkinter.

* Additions to UI, still non-functional

* UI now functional, untested

* UI now functional, untested

* Added saving configs

* Saving and loading now functional

* Fixed sliders not loading

* Cleaned up duplicate arrays

* Cleaned up duplicate arrays

* Fixed loading bugs

* wip fixing all the broken parameters. PLEASE test before you commit

* further cleaning

* bugfix completed for gui. now evaluating save and load

* cleanup prepare to merge

---------

Co-authored-by: Concedo <39025047+LostRuins@users.noreply.github.com>
phuongncn pushed a commit to phuongncn/llama.cpp-gx10-dgx-sparks-deepseekv4 that referenced this pull request Apr 28, 2026
…gml-org#284)

* llama-bench: enable having different number of threads for tg and pp

* Add -tgb to usage

---------

Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
bong-water-water-bong added a commit to 1bit-MONSTER/llama.cpp that referenced this pull request Oct 5, 2026
ggml-hrx: the f32_f32 WMMA cores accumulate in f32 (engine ggml-org#284)
Copilot AI pushed a commit to imaami/llama.cpp that referenced this pull request Oct 6, 2026
…gml-org#284)

* dflash: add beam selection and reuse Metal verify data

* metal: few-row tensor tile for PQ2_0 at speculative verify widths (ggml-org#286)

Adds kernel_mul_mm_pq2_0_f32_fewrow, a register-resident matmul2d tile of
16 src1 rows by 32 weight rows per simdgroup. PQ2_0 weights are decoded
straight into the right-hand cooperative tensor, so the K loop needs no
threadgroup staging and no barriers. On Metal 4 devices the generic tensor
mul_mm tile is 64 x 128, which computes far more columns than a 4..16 row
verify batch needs, and mul_mv_ext is ALU-bound at 2..8 columns.

The route is taken only when the device has the tensor API, src0 is PQ2_0,
src1 is contiguous F32, there is no batching in dims 2/3, and ne11 is in
[GGML_METAL_PQ2_0_FEWROW_MIN, _MAX] (default 3..32). Below 3 columns the
existing mat-vec paths stay faster. GGML_METAL_PQ2_0_FEWROW=0 turns the
route off, and GGML_METAL_PQ2_0_FEWROW_CFG selects the column-block by
K-split layout.

test-backend-ops perf on M5 Pro, 17408 x 5120, best of 2 (us/run):
  n=3   187.0 -> 151.3   (3-column mul_mv from the parent commit)
  n=8   484.5 -> 157.4   (mul_mv_ext)
  n=16  855.9 -> 172.8   (tensor mul_mm)
  n=32  864.0 -> 328.3
n=1..2 are slower on the new tile (150 vs 71 and 130 us), hence min 3.

End to end on M5 Pro (27B PQ2_0 target, DFlash2 drafter, greedy, depth 7,
8 prompts x 3 reps): with this route off, speculation runs at 0.73x of
plain decoding; with it on, 1.57x (BF16 drafter) to 1.64x (Q8_0 drafter).
Output matches plain decoding except one near-tie (top-2 logprob gap
0.00047) where the half-precision activations pick the other token.

The cooperative tensor element layout was probed on M5 with
get_multidimensional_index before writing the decode.

test-backend-ops gains PQ2_0 MUL_MAT cases for every n from 1 to 32 plus
64 and 512, and perf cases at the 27B projection shapes, plus flash attention perf cases
at the DFlash2 drafter and 27B verify shapes.

The few-row approach is adapted from the qmm_m16_block core in the Yukon
MLX.fast promoted submissions (polymorf and later contributors), used
under Apache-2.0. The PQ2_0 block layout differs, so the decode and
dispatch are new code.

(cherry picked from commit 3285e75)
Assisted-by: GitHub Copilot
Copilot AI pushed a commit to imaami/llama.cpp that referenced this pull request Oct 6, 2026
…gml-org#284)

* dflash: add beam selection and reuse Metal verify data

* metal: few-row tensor tile for PQ2_0 at speculative verify widths (ggml-org#286)

Adds kernel_mul_mm_pq2_0_f32_fewrow, a register-resident matmul2d tile of
16 src1 rows by 32 weight rows per simdgroup. PQ2_0 weights are decoded
straight into the right-hand cooperative tensor, so the K loop needs no
threadgroup staging and no barriers. On Metal 4 devices the generic tensor
mul_mm tile is 64 x 128, which computes far more columns than a 4..16 row
verify batch needs, and mul_mv_ext is ALU-bound at 2..8 columns.

The route is taken only when the device has the tensor API, src0 is PQ2_0,
src1 is contiguous F32, there is no batching in dims 2/3, and ne11 is in
[GGML_METAL_PQ2_0_FEWROW_MIN, _MAX] (default 3..32). Below 3 columns the
existing mat-vec paths stay faster. GGML_METAL_PQ2_0_FEWROW=0 turns the
route off, and GGML_METAL_PQ2_0_FEWROW_CFG selects the column-block by
K-split layout.

test-backend-ops perf on M5 Pro, 17408 x 5120, best of 2 (us/run):
  n=3   187.0 -> 151.3   (3-column mul_mv from the parent commit)
  n=8   484.5 -> 157.4   (mul_mv_ext)
  n=16  855.9 -> 172.8   (tensor mul_mm)
  n=32  864.0 -> 328.3
n=1..2 are slower on the new tile (150 vs 71 and 130 us), hence min 3.

End to end on M5 Pro (27B PQ2_0 target, DFlash2 drafter, greedy, depth 7,
8 prompts x 3 reps): with this route off, speculation runs at 0.73x of
plain decoding; with it on, 1.57x (BF16 drafter) to 1.64x (Q8_0 drafter).
Output matches plain decoding except one near-tie (top-2 logprob gap
0.00047) where the half-precision activations pick the other token.

The cooperative tensor element layout was probed on M5 with
get_multidimensional_index before writing the decode.

test-backend-ops gains PQ2_0 MUL_MAT cases for every n from 1 to 32 plus
64 and 512, and perf cases at the 27B projection shapes, plus flash attention perf cases
at the DFlash2 drafter and 27B verify shapes.

The few-row approach is adapted from the qmm_m16_block core in the Yukon
MLX.fast promoted submissions (polymorf and later contributors), used
under Apache-2.0. The PQ2_0 block layout differs, so the decode and
dispatch are new code.

(cherry picked from commit 3285e75)
Assisted-by: GitHub Copilot
imaami pushed a commit to imaami/llama.cpp that referenced this pull request Oct 6, 2026
…gml-org#284)

* dflash: add beam selection and reuse Metal verify data

* metal: few-row tensor tile for PQ2_0 at speculative verify widths (ggml-org#286)

Adds kernel_mul_mm_pq2_0_f32_fewrow, a register-resident matmul2d tile of
16 src1 rows by 32 weight rows per simdgroup. PQ2_0 weights are decoded
straight into the right-hand cooperative tensor, so the K loop needs no
threadgroup staging and no barriers. On Metal 4 devices the generic tensor
mul_mm tile is 64 x 128, which computes far more columns than a 4..16 row
verify batch needs, and mul_mv_ext is ALU-bound at 2..8 columns.

The route is taken only when the device has the tensor API, src0 is PQ2_0,
src1 is contiguous F32, there is no batching in dims 2/3, and ne11 is in
[GGML_METAL_PQ2_0_FEWROW_MIN, _MAX] (default 3..32). Below 3 columns the
existing mat-vec paths stay faster. GGML_METAL_PQ2_0_FEWROW=0 turns the
route off, and GGML_METAL_PQ2_0_FEWROW_CFG selects the column-block by
K-split layout.

test-backend-ops perf on M5 Pro, 17408 x 5120, best of 2 (us/run):
  n=3   187.0 -> 151.3   (3-column mul_mv from the parent commit)
  n=8   484.5 -> 157.4   (mul_mv_ext)
  n=16  855.9 -> 172.8   (tensor mul_mm)
  n=32  864.0 -> 328.3
n=1..2 are slower on the new tile (150 vs 71 and 130 us), hence min 3.

End to end on M5 Pro (27B PQ2_0 target, DFlash2 drafter, greedy, depth 7,
8 prompts x 3 reps): with this route off, speculation runs at 0.73x of
plain decoding; with it on, 1.57x (BF16 drafter) to 1.64x (Q8_0 drafter).
Output matches plain decoding except one near-tie (top-2 logprob gap
0.00047) where the half-precision activations pick the other token.

The cooperative tensor element layout was probed on M5 with
get_multidimensional_index before writing the decode.

test-backend-ops gains PQ2_0 MUL_MAT cases for every n from 1 to 32 plus
64 and 512, and perf cases at the 27B projection shapes, plus flash attention perf cases
at the DFlash2 drafter and 27B verify shapes.

The few-row approach is adapted from the qmm_m16_block core in the Yukon
MLX.fast promoted submissions (polymorf and later contributors), used
under Apache-2.0. The PQ2_0 block layout differs, so the decode and
dispatch are new code.

(cherry picked from commit 3285e75)
Assisted-by: GitHub Copilot
imaami pushed a commit to imaami/llama.cpp that referenced this pull request Oct 6, 2026
…gml-org#284)

* dflash: add beam selection and reuse Metal verify data

* metal: few-row tensor tile for PQ2_0 at speculative verify widths (ggml-org#286)

Adds kernel_mul_mm_pq2_0_f32_fewrow, a register-resident matmul2d tile of
16 src1 rows by 32 weight rows per simdgroup. PQ2_0 weights are decoded
straight into the right-hand cooperative tensor, so the K loop needs no
threadgroup staging and no barriers. On Metal 4 devices the generic tensor
mul_mm tile is 64 x 128, which computes far more columns than a 4..16 row
verify batch needs, and mul_mv_ext is ALU-bound at 2..8 columns.

The route is taken only when the device has the tensor API, src0 is PQ2_0,
src1 is contiguous F32, there is no batching in dims 2/3, and ne11 is in
[GGML_METAL_PQ2_0_FEWROW_MIN, _MAX] (default 3..32). Below 3 columns the
existing mat-vec paths stay faster. GGML_METAL_PQ2_0_FEWROW=0 turns the
route off, and GGML_METAL_PQ2_0_FEWROW_CFG selects the column-block by
K-split layout.

test-backend-ops perf on M5 Pro, 17408 x 5120, best of 2 (us/run):
  n=3   187.0 -> 151.3   (3-column mul_mv from the parent commit)
  n=8   484.5 -> 157.4   (mul_mv_ext)
  n=16  855.9 -> 172.8   (tensor mul_mm)
  n=32  864.0 -> 328.3
n=1..2 are slower on the new tile (150 vs 71 and 130 us), hence min 3.

End to end on M5 Pro (27B PQ2_0 target, DFlash2 drafter, greedy, depth 7,
8 prompts x 3 reps): with this route off, speculation runs at 0.73x of
plain decoding; with it on, 1.57x (BF16 drafter) to 1.64x (Q8_0 drafter).
Output matches plain decoding except one near-tie (top-2 logprob gap
0.00047) where the half-precision activations pick the other token.

The cooperative tensor element layout was probed on M5 with
get_multidimensional_index before writing the decode.

test-backend-ops gains PQ2_0 MUL_MAT cases for every n from 1 to 32 plus
64 and 512, and perf cases at the 27B projection shapes, plus flash attention perf cases
at the DFlash2 drafter and 27B verify shapes.

The few-row approach is adapted from the qmm_m16_block core in the Yukon
MLX.fast promoted submissions (polymorf and later contributors), used
under Apache-2.0. The PQ2_0 block layout differs, so the decode and
dispatch are new code.

(cherry picked from commit 3285e75)
Assisted-by: GitHub Copilot
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants