Skip to content

AVX2: Speed up large batch size prompt processing of IQ models - #27402

Merged
bartowski1182 merged 13 commits into
ggml-org:masterfrom
bartowski1182:quant-run-speed
Aug 31, 2026
Merged

bartowski1182 merged 13 commits into
ggml-org:masterfrom
bartowski1182:quant-run-speed

Conversation

@bartowski1182

@bartowski1182 bartowski1182 commented Aug 19, 2026 •

Copy link
Copy Markdown
Contributor

Overview

IQ quants are particularly slow on CPU at large batch sizes (what you'd see for imatrix and perplexity)

Part of this comes from the fact that during a 512-token batch, every weight in the model is decoded 512 times from the associated lookup table

This change introduces a GEMM panel, block_iqp_x8, which is 8 weight rows x 256 columns:

struct block_iqp_x8 {
    float   dfac[8];               // per row: the float part of the scale (d * 2^-k)
    int32_t bias[8];               // per row: unsigned-activation correction
    int8_t  iscales[16 * 8];       // per (sub-block, row): the integer part of the scale
    int8_t  qs[256 * 8];           // the decoded weights, interleaved
};

Now, instead of decoding each weight 512 times at batch 512, we decode 8 rows at a time into a cache-sized int8 tile and run integer GEMM over it.

For dense models, this is a huge increase, up to 10x

For MoE it's more modest, more like 2x, but still quite good

PPL shifts a tiny bit, only ~0.24%, aka margin of error

Design details from Claude

Four design details make it work:

  1. Interleaving 8 rows. qs[sb128 + g32 + row4 + k] holds column sb16 + g*4 + k. So one 32-byte load pulls 4 consecutive columns for each of 8 rows — exactly the operand shape _mm256_dpbusd_epi32 wants against a broadcast int32 of 4 activations. Eight output rows advance per instruction.

  2. Sub-blocks of 16, not 32. iq2_xs/iq2_s take a group of 32 weights' scale from two nibbles, one per half. One scale per 32 couldn't represent them exactly, so 16 is used uniformly for all eight types → one kernel.

  3. Split scale = float × integer (this is commit 4f7e113). Every supported type's sub-block scale factors as (d · 2⁻ᵏ) · small_int — e.g. iq2_xs is d·(0.5+ls)·0.25 = (d/8)·(2ls+1). So dfac holds the float, iscales the integer. The kernel then applies sub-block scales in integer (mullo_epi32), accumulates the whole 256-weight super-block exactly in int32, and touches float exactly once per (row, super-block). That's 16 float roundings → 1. A float64 reference harness measured the result as 12–22% more accurate than upstream vec_dot (the previous float-scale version had been 2.6× less accurate).

  4. The bias trick. VNNI's dpbusd needs unsigned × signed, but both operands here are signed. Instead of fixing that with _mm256_sign_epi8 per operand, activations are fed as y ^ 0x80 (i.e. y + 128), and the resulting spurious 128·Σw term — a per-row, per-super-block constant — is precomputed into bias at decode time and subtracted with one int32 op. Non-VNNI AVX2 builds fall back to the sign trick with bias disabled, because the maddubs int16 accumulator would overflow with unsigned operands.

Additional information

On dense models, this only kicks in when batch >= 32 8 columns, less than that was found to have lower performance, set with GGML_IQP_MIN_BATCH in ggml/src/ggml-cpu/iqp.h (edit: lowered to 8 with a better kernel)

for MoE, it instead gates per-expert, and only works when needs to see 16 8 columns to route through the new panel, otherwise old performance is better, set with GGML_IQP_MIN_BATCH_ID in ggml/src/ggml-cpu/iqp.h. (edit: also lowered to 8 with better kernel)

Can be disabled with GGML_NO_IQ_PANEL=1

Confirmed with tests that setting GGML_NO_IQ_PANEL=1 returns original performance and PPL completely, as well as when using a lower batch size (-ub 16). At -ub 16, there is still an ~80% speed up in performance on dense

Benchmark numbers

I ran PPL against master and this PR to get speed and numbers on --chunks 50 for Qwen3.6-27B and Qwen3.6-35B-A3B on EPYC 9654 using 24 threads

Created pure IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, and IQ4_XS. Made pure to make sure each tensor type is fully exercised.

Updated figures:

image

Speed:

Model Quant ub8 ub32 ub128 ub512
Qwen3.6-27B iq1_m 12.08 → 17.16 (1.42×) 13.33 → 39.50 (2.96×) 14.23 → 62.54 (4.39×) 14.36 → 73.70 (5.13×)
iq1_s 11.98 → 18.11 (1.51×) 13.31 → 41.18 (3.09×) 13.84 → 63.67 (4.60×) 14.01 → 73.60 (5.25×)
iq2_s 12.07 → 20.98 (1.74×) 13.43 → 45.84 (3.41×) 13.98 → 65.58 (4.69×) 14.13 → 74.45 (5.27×)
iq2_xs 12.25 → 17.08 (1.39×) 13.73 → 40.04 (2.92×) 14.27 → 62.53 (4.38×) 14.42 → 73.70 (5.11×)
iq2_xxs 11.66 → 16.45 (1.41×) 12.89 → 39.19 (3.04×) 13.40 → 61.82 (4.61×) 13.51 → 73.47 (5.44×)
iq3_s 7.95 → 17.26 (2.17×) 8.51 → 41.66 (4.90×) 8.72 → 63.39 (7.27×) 8.77 → 73.84 (8.42×)
iq3_xxs 9.07 → 15.32 (1.69×) 9.87 → 38.08 (3.86×) 10.17 → 60.93 (5.99×) 10.25 → 73.52 (7.17×)
iq4_xs 17.75 → 25.46 (1.43×) 20.85 → 49.72 (2.38×) 22.01 → 67.62 (3.07×) 22.42 → 75.99 (3.39×)
Qwen3.6-35B-A3B iq1_m 65.99 → 70.38 (1.07×) 94.16 → 133.51 (1.42×) 106.81 → 208.87 (1.96×) 111.09 → 270.72 (2.44×)
iq1_s 65.77 → 71.03 (1.08×) 93.37 → 136.41 (1.46×) 106.14 → 210.16 (1.98×) 109.49 → 271.44 (2.48×)
iq2_s 66.56 → 72.96 (1.10×) 93.49 → 138.01 (1.48×) 105.68 → 214.45 (2.03×) 109.23 → 273.43 (2.50×)
iq2_xs 67.24 → 70.61 (1.05×) 94.35 → 134.41 (1.42×) 106.47 → 206.45 (1.94×) 109.96 → 270.01 (2.46×)
iq2_xxs 64.39 → 69.67 (1.08×) 88.48 → 133.62 (1.51×) 102.37 → 206.35 (2.02×) 105.68 → 273.43 (2.59×)
iq3_s 51.15 → 60.38 (1.18×) 66.57 → 115.25 (1.73×) 72.59 → 185.76 (2.56×) 74.27 → 262.56 (3.54×)
iq3_xxs 56.51 → 61.12 (1.08×) 74.10 → 118.47 (1.60×) 82.70 → 188.50 (2.28×) 85.42 → 262.06 (3.07×)
iq4_xs 73.94 → 77.26 (1.04×) 120.35 → 152.68 (1.27×) 142.17 → 236.08 (1.66×) 154.27 → 293.41 (1.90×)

Perplexity:

Model Quant ub8 ub32 ub128 ub512
Qwen3.6-27B iq1_m 13.5217 → 13.5001 (-0.16%) 14.1851 → 14.1505 (-0.24%) 12.8995 → 12.8299 (-0.54%) 12.8952 → 12.8573 (-0.29%)
iq1_s 18.8895 → 18.8092 (-0.43%) 19.0357 → 19.0554 (+0.10%) 17.2344 → 17.1928 (-0.24%) 17.1661 → 17.1822 (+0.09%)
iq2_s 8.1025 → 8.1011 (-0.02%) 8.4542 → 8.4466 (-0.09%) 7.7211 → 7.7219 (+0.01%) 7.7392 → 7.6964 (-0.55%)
iq2_xs 8.9372 → 8.9305 (-0.07%) 9.2052 → 9.2050 (-0.00%) 8.4168 → 8.4295 (+0.15%) 8.4474 → 8.3950 (-0.62%)
iq2_xxs 9.6466 → 9.6452 (-0.01%) 10.1724 → 10.1607 (-0.12%) 9.2808 → 9.2768 (-0.04%) 9.3019 → 9.2849 (-0.18%)
iq3_s 7.0802 → 7.0787 (-0.02%) 7.4442 → 7.4469 (+0.04%) 6.6514 → 6.6520 (+0.01%) 6.6336 → 6.6398 (+0.09%)
iq3_xxs 7.3029 → 7.2900 (-0.18%) 7.6893 → 7.6891 (-0.00%) 6.9429 → 6.9441 (+0.02%) 6.9333 → 6.9408 (+0.11%)
iq4_xs 6.9964 → 6.9766 (-0.28%) 7.3233 → 7.3036 (-0.27%) 6.6012 → 6.5835 (-0.27%) 6.6033 → 6.5726 (-0.46%)
Qwen3.6-35B-A3B iq1_m 14.7116 → 14.7737 (+0.42%) 16.0681 → 16.0237 (-0.28%) 14.4098 → 14.2752 (-0.93%) 14.3766 → 14.3442 (-0.23%)
iq1_s 24.2457 → 24.3448 (+0.41%) 26.9374 → 26.9539 (+0.06%) 23.0651 → 22.9926 (-0.31%) 23.3641 → 23.1182 (-1.05%)
iq2_s 8.5473 → 8.5412 (-0.07%) 9.4422 → 9.4887 (+0.49%) 8.3438 → 8.4120 (+0.82%) 8.3330 → 8.3738 (+0.49%)
iq2_xs 9.0943 → 9.0833 (-0.12%) 10.1651 → 10.1508 (-0.14%) 8.9008 → 8.8918 (-0.10%) 8.8810 → 8.8749 (-0.07%)
iq2_xxs 10.9124 → 10.8715 (-0.37%) 12.3101 → 12.2612 (-0.40%) 10.9847 → 10.9377 (-0.43%) 11.0131 → 10.8834 (-1.18%)
iq3_s 7.0690 → 7.0806 (+0.16%) 7.6931 → 7.6639 (-0.38%) 6.7341 → 6.7085 (-0.38%) 6.7541 → 6.7269 (-0.40%)
iq3_xxs 7.2852 → 7.2931 (+0.11%) 7.9733 → 7.9480 (-0.32%) 6.9312 → 6.9747 (+0.63%) 6.9089 → 6.9594 (+0.73%)
iq4_xs 6.7495 → 6.7288 (-0.31%) 7.2500 → 7.2000 (-0.69%) 6.3037 → 6.2672 (-0.58%) 6.2625 → 6.2554 (-0.11%)

I ran at various chunk lengths for each ubatch to save time, all were run with these settings:

ubatch 512, 8 chunks
ubatch 128, 8 chunks
ubatch 32, 12 chunks
ubatch 8, 20 chunks

I also tested:

./build/bin/test-backend-ops test -b CPU -o MUL_MAT
./build/bin/test-backend-ops test -b CPU -o MUL_MAT_ID

On a Ryzen 9 7950X3D and an Intel Core Ultra 7 358H since they have VNNI support, both gave OK

On the Ryzen 9 I got these numbers:

== Qwen3.6-27B-pure-iq2_xs [default]
   PPL = 9.1726 +/- 0.46416   eval 113.46s (45.13 tok/s)   116s wall
== Qwen3.6-27B-pure-iq2_xs [panel-off]
   PPL = 9.2220 +/- 0.46841   eval 495.97s (10.32 tok/s)   499s wall
== Qwen3.6-27B-pure-iq2_xs [ub8] -ub 8
   PPL = 9.1662 +/- 0.46399   eval 465.61s (11.00 tok/s)   468s wall
(forgot to measure panel-off of ub 8 but panel on is faster than 512 soo...)
== Qwen3.6-35B-A3B-pure-iq3_s [default]
   PPL = 7.6474 +/- 0.37486   eval 26.56s (192.77 tok/s)   29s wall
== Qwen3.6-35B-A3B-pure-iq3_s [panel-off]
   PPL = 7.6666 +/- 0.37649   eval 95.13s (53.82 tok/s)   97s wall
== Qwen3.6-35B-A3B-pure-iq3_s [ub8] -ub 8
   PPL = 7.6881 +/- 0.37842   eval 112.05s (45.69 tok/s)   115s wall
== Qwen3.6-35B-A3B-pure-iq3_s [ub8-panel-off] -ub 8
   PPL = 7.6440 +/- 0.37451   eval 132.23s (38.72 tok/s)   134s wall
original speeds/ppl just for posterity

These are the most extremely differences because it's at a big batch size (512), lower batch sizes get smaller increases

Model PPL master PPL PR PPL diff tok/s master tok/s PR tok/s diff
Qwen3.6-27B-pure-iq1_m 12.1242 +/- 0.27911 12.1355 +/- 0.27961 +0.0113 (+0.09%) 9.10 69.59 +60.49 (+664.7%)
Qwen3.6-27B-pure-iq1_s 17.1841 +/- 0.41605 17.2043 +/- 0.41636 +0.0202 (+0.12%) 8.57 70.10 +61.53 (+718.0%)
Qwen3.6-27B-pure-iq2_s 7.4571 +/- 0.16908 7.4440 +/- 0.16864 -0.0131 (-0.18%) 7.62 67.81 +60.19 (+789.9%)
Qwen3.6-27B-pure-iq2_xs 8.0930 +/- 0.18622 8.0798 +/- 0.18562 -0.0132 (-0.16%) 8.78 67.82 +59.04 (+672.4%)
Qwen3.6-27B-pure-iq2_xxs 8.5515 +/- 0.19470 8.5466 +/- 0.19442 -0.0049 (-0.06%) 7.21 68.19 +60.98 (+845.8%)
Qwen3.6-27B-pure-iq3_s 6.4753 +/- 0.14089 6.4779 +/- 0.14108 +0.0026 (+0.04%) 4.75 65.45 +60.70 (+1277.9%)
Qwen3.6-27B-pure-iq3_xxs 6.6138 +/- 0.14414 6.6223 +/- 0.14448 +0.0085 (+0.13%) 6.12 67.43 +61.31 (+1001.8%)
Qwen3.6-27B-pure-iq4_xs 6.4100 +/- 0.14195 6.4073 +/- 0.14187 -0.0027 (-0.04%) 22.07 69.19 +47.12 (+213.5%)
Qwen3.6-35B-A3B-pure-iq1_m 12.9822 +/- 0.31998 13.0037 +/- 0.32059 +0.0215 (+0.17%) 111.28 244.04 +132.76 (+119.3%)
Qwen3.6-35B-A3B-pure-iq1_s 20.5812 +/- 0.56679 20.5967 +/- 0.56756 +0.0155 (+0.08%) 110.13 245.92 +135.79 (+123.3%)
Qwen3.6-35B-A3B-pure-iq2_s 7.5883 +/- 0.16738 7.5798 +/- 0.16713 -0.0085 (-0.11%) 111.48 229.51 +118.03 (+105.9%)
Qwen3.6-35B-A3B-pure-iq2_xs 8.1627 +/- 0.18140 8.1432 +/- 0.18089 -0.0195 (-0.24%) 110.29 234.09 +123.80 (+112.2%)
Qwen3.6-35B-A3B-pure-iq2_xxs 9.9025 +/- 0.22833 9.8890 +/- 0.22815 -0.0135 (-0.14%) 106.57 231.40 +124.83 (+117.1%)
Qwen3.6-35B-A3B-pure-iq3_s 6.4325 +/- 0.13703 6.4316 +/- 0.13695 -0.0009 (-0.01%) 74.47 205.42 +130.95 (+175.8%)
Qwen3.6-35B-A3B-pure-iq3_xxs 6.5745 +/- 0.14131 6.5797 +/- 0.14136 +0.0052 (+0.08%) 85.61 221.24 +135.63 (+158.4%)
Qwen3.6-35B-A3B-pure-iq4_xs 6.1650 +/- 0.13255 6.1633 +/- 0.13263 -0.0017 (-0.03%) 156.09 245.02 +88.93 (+57.0%)

Note, since some of these are extremely long running even at only 50 chunks, the performance numbers may vary slightly, but the gains were seen repeatedly.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, exclusively, my idea was to repack the weights and use up extra RAM which was a useful PoC, Claude found the panelling and a way to avoid extra usage. The implementation is entirely AI, validated only by testing.

@github-actions github-actions Bot added the ggml changes relating to the ggml tensor library for machine learning label Aug 19, 2026
@bartowski1182
bartowski1182 marked this pull request as ready for review August 22, 2026 00:52
@bartowski1182

Copy link
Copy Markdown
Contributor Author

Updated with latest numbers for performance, looks quite good across the board.

Also noted that I ran the test ops on machines with VNNI for completeness

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, this is not really repacking of the weights - they stay the same, we just add some on-the-fly repacking to the work buffers. If this is correct, then the repack.cpp should not be modified. Most likely you need to reorganize the code similar to llamafile and not modify the repack logic at all.

@bartowski1182

Copy link
Copy Markdown
Contributor Author

@ggerganov it's not exactly repacking but it is doing similar work to other stuff already in repack.cpp, namely the row interleaving

I guess the question is, does repack only apply to load-time interleaving or also transient interleaving?

If it makes more sense to move it to avoid calling it something it isn't, I can do that, will need to duplicate some existing work around architecture mapping but it's definitely doable

@ggerganov

Copy link
Copy Markdown
Member

I guess the question is, does repack only apply to load-time interleaving or also transient interleaving?

Repacking is just for load-time. It uses an extra buffer type and modifies the contents of the weight tensors.

Runtime-only optimizations like in this PR should be implemented separately similar to llamafile and simd-sgemm.

@bartowski1182

Copy link
Copy Markdown
Contributor Author

@ggerganov okay it's all now self contained in the iqp.cpp file, repack is now untouched, running sanity checks to make sure nothing got dropped (mostly around disabling when it shouldn't be used), the performance looks identical

@bartowski1182

bartowski1182 commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor Author

@ggerganov let me know if you need anything else from me (more explanation, more tests, etc)

all sanity checks look good

@ggerganov

Copy link
Copy Markdown
Member

Please rebase on latest master.

@ggerganov ggerganov self-assigned this Aug 25, 2026
@bartowski1182
bartowski1182 force-pushed the quant-run-speed branch 2 times, most recently from 809f6da to 9f6857f Compare August 25, 2026 17:19
@bartowski1182

Copy link
Copy Markdown
Contributor Author

Rebased and squashed

@ggerganov

Copy link
Copy Markdown
Member

Add yourself to the CODEOWNERS for the new iqp sources. And just in case, run a thread-sanitizer runs to make sure there are no data races in the new logic. Would be nice to get a second validation of the results - I don't have a suitable machine to test atm.

@bartowski1182

Copy link
Copy Markdown
Contributor Author

Okay here's what I ran, let me know if it was right:

cmake -B build-thread-sani \
    -DCMAKE_BUILD_TYPE=RelWithDebInfo \
    -DLLAMA_SANITIZE_THREAD=ON \
    -DGGML_OPENMP=OFF -DLLAMA_BUILD_UI=OFF -DLLAMA_USE_PREBUILT_UI=OFF

cmake --build build-thread-sani -j

setarch "$(uname -m)" -R ./build-thread-sani/bin/test-backend-ops test -b CPU -o MUL_MAT
Testing 1 devices

Backend 1/1: CPU
  Device description: AMD Eng Sample: 100-000000894-04
  Device memory: 773295 MB (773295 MB free)

  MUL_MAT(type_a=f32,type_b=f32,m=16,n=1,k=256,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): OK
...
  MUL_MAT(type_a=f32,type_b=f32,m=129,n=1,k=1057,bs=[8,3],nr=[4,1],per=[0,1,2,3],k_v=2113,o=1,src_overlap=0): OK
  1204/1204 tests passed
  Backend CPU: OK
1/1 backends passed
OK

setarch "$(uname -m)" -R ./build-thread-sani/bin/test-backend-ops test -b CPU -o MUL_MAT_ID
Testing 1 devices

Backend 1/1: CPU
  Device description: AMD Eng Sample: 100-000000894-04
  Device memory: 773295 MB (773295 MB free)

  MUL_MAT_ID(type_a=f16,type_b=f32,n_mats=16,n_used=16,b=0,m=32,n=1024,k=16): OK
...
  MUL_MAT_ID(type_a=bf16,type_b=f32,n_mats=4,n_used=2,b=0,m=512,n=32,k=256): OK
  872/872 tests passed
  Backend CPU: OK
1/1 backends passed
OK

@bartowski1182

Copy link
Copy Markdown
Contributor Author

I have a few other CPUs I can test on and also can ask people to help test if you need more confirmation, just say the word :)

@ggerganov

Copy link
Copy Markdown
Member

You will also need GGML_SANITIZE_THREAD in addition to LLAMA_SANITIZE_THREAD.

Also, add yourself to the CODEOWNERS for the new iqp sources.

@bartowski1182

Copy link
Copy Markdown
Contributor Author

Okay I'll try to run that when I get the chance!

I did add myself to the CODEOWNERS file, did I do it right?

42a8aba

@bartowski1182

Copy link
Copy Markdown
Contributor Author

So I ran

cmake -B build-thread-sani  -DCMAKE_BUILD_TYPE=RelWithDebInfo  -DLLAMA_SANITIZE_THREAD=ON  -DGGML_OPENMP=OFF -DLLAMA_BUILD_UI=OFF -DLLAMA_USE_PREBUILT_UI=OFF -DGGML_SANITIZE_THREAD=ON

and

setarch "$(uname -m)" -R ./build-thread-sani/bin/test-backend-ops test -b CPU -o MUL_MAT

it returned all OK, but it ran on all my cores for about 19 hours... Do I need to also run MUL_MAT_ID? haha..

@bartowski1182

Copy link
Copy Markdown
Contributor Author

@ggerganov if you need me to run the other i'll do it overnight tonight, not sure what the difference is between them but i'm guessing the GGML_SANITIZE_THREADS makes it way slower to run

Comment thread ggml/src/ggml-cpu/ggml-cpu.c
Comment thread ggml/src/ggml-cpu/iqp.h Outdated
@bartowski1182

Copy link
Copy Markdown
Contributor Author

also rebased on master again, hence the force-push

Comment thread ggml/src/ggml-cpu/iqp.h Outdated
Comment thread ggml/src/ggml-cpu/iqp.h Outdated
Comment thread ggml/src/ggml-cpu/iqp.h Outdated
Comment thread ggml/src/ggml-cpu/iqp.h Outdated
Comment thread ggml/src/ggml-cpu/iqp.h Outdated
Comment thread ggml/src/ggml-cpu/iqp.h Outdated
Comment thread ggml/src/ggml-cpu/ggml-cpu.c
Comment thread ggml/src/ggml-common.h
@@ -1131,7 +1131,7 @@ GGML_TABLE_END()
#define NGRID_IQ1S 2048
#define IQ1S_DELTA 0.125f
#define IQ1M_DELTA 0.125f
#if defined(GGML_COMMON_IMPL_C)
#if defined(GGML_COMMON_IMPL_C) || defined(GGML_COMMON_IMPL_CPP)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

noticed that technically this change removes the iq1s_grid_gpu from any area that has GGML_COMMON_IMPL_CPP, however obviously none of them actually use it currently, just wanted to call attention to it in case it raises any questions about the grid that they have access to switching from this change

@bartowski1182

Copy link
Copy Markdown
Contributor Author

Moved IQP mul_mat_id test and rebased on latest

@bartowski1182
bartowski1182 merged commit 85c5522 into ggml-org:master Aug 31, 2026
26 of 28 checks passed
@bartowski1182

Copy link
Copy Markdown
Contributor Author

@CISC I just woke up so I trusted Claude to investigate so take this with a grain of salt, but it seems to think it's actually a CI ccache issue, not a bug in the code:

Your PR's one-line edit to ggml/src/ggml-common.h invalidated cache entries only for translation units that include that header: ggml-cpu.cpp includes it via repack.h, and amx/mmq.cpp via ggml-quants.h — but amx/amx.cpp does not include ggml-common.h at all, so its stale object was reused

Fix is allegedly to clear the ccache for those jobs and re run, or add -DGGML_NATIVE=OFF to those two failing jobs

@CISC

CISC commented Sep 1, 2026

Copy link
Copy Markdown
Member

@CISC I just woke up so I trusted Claude to investigate so take this with a grain of salt, but it seems to think it's actually a CI ccache issue, not a bug in the code:

Yes, most likely, it's a bit weird, will try deleting caches.

@CISC

CISC commented Sep 1, 2026

Copy link
Copy Markdown
Member

@CISC I just woke up so I trusted Claude to investigate so take this with a grain of salt, but it seems to think it's actually a CI ccache issue, not a bug in the code:

Yes, most likely, it's a bit weird, will try deleting caches.

Seems to have worked:
https://github.com/ggml-org/llama.cpp/actions/runs/33489948551/job/99798817425

@bartowski1182

Copy link
Copy Markdown
Contributor Author

Great thanks for handling it!

ilmmatias pushed a commit to ilmmatias/llama.cpp that referenced this pull request Sep 1, 2026
…org#27402)

* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
@jbooth

jbooth commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Hi, I just coincidentally had the same idea and implemented something similar for k-quants. I put a proposal for merging the 2 approaches into a single code path on my PR here: #27851 @bartowski1182 if you want to take a look.

OllyJohnston added a commit to OllyJohnston/llama.cpp that referenced this pull request Sep 2, 2026
Port upstream PR ggml-org#27402 (grid IQ batched GEMM) onto the
bmoe/expert-ready-hook line:

- new iqp.cpp/iqp.h: decode 8 src0 rows at a time into per-thread
  int8 panels (block_iqp_x8) and run an integer gemm against all
  src1 columns, for the grid IQ types (IQ1_S/M, IQ2_XXS/XS/S,
  IQ3_XXS/S, IQ4_XS/NL) - up to 8-10x faster CPU prefill at large
  batch sizes, ~2x on CPU MoE layers
- ggml-cpu.c: dispatch mul_mat and mul_mat_id to the iqp path after
  the src1->q8_K barrier (mirrors the upstream placement: the iqp
  path consumes the q8_K rows from the work buffer); the per-expert
  mul_mat_id dispatch comes after the expert-ready hook so the
  streamer's residency block still fires on both paths
- graph_plan reserves one scratch panel per thread for both ops
- ggml-common.h: widen iq1s_grid table guard to C++ translation units
- tests: 10-row-batch IQP coverage on all grid IQ types for MUL_MAT
  and MUL_MAT_ID

Verified: test-backend-ops 883/883 MUL_MAT + MUL_MAT_ID on CPU
(x86 AVX2) including the new 10-row batch cases; CUDA0 topk/fusion
416/416 unchanged; bmoe byte-identity gates 13/16 (pre-existing
streaming-init failures only).
aukarande added a commit to aukarande/llama.cpp that referenced this pull request Sep 3, 2026
@BrewTestBot BrewTestBot mentioned this pull request Sep 4, 2026
1 task done
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
…org#27402)

* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…org#27402)

* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…org#27402)

* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
…org#27402)

* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Sep 29, 2026
…org#27402)

* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Sep 29, 2026
…org#27402)

* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants