Skip to content

B1 mtp qwen rebase - #11

Merged
Ooooze merged 11 commits into
feature/turboquant-kv-cachefrom
b1-mtp-qwen-rebase
May 12, 2026
Merged

Ooooze merged 11 commits into
feature/turboquant-kv-cachefrom
b1-mtp-qwen-rebase

Conversation

@Ooooze

@Ooooze Ooooze commented May 12, 2026

Copy link
Copy Markdown

Overview

Additional information

Requirements

am17an and others added 11 commits May 8, 2026 14:30
Currently speculative checkpoint needs to restart from a checkpoint
after some draft tokens are not accepted, this leads to some wastage in
running the target again. This PR adds the ability to rollback upto
`draft_max` by storing the GDN intermediates.
Recovery snapshot from agent transcript replay after accidental
`git checkout` of working-tree-only NextN changes. Build is clean,
but NextN inference is broken: target argmax produces garbage tokens
when `cparams.embeddings_pre_norm=true` is paired with the server-side
`nextn_prefill_all_outputs` + `do_checkpoint=false` overrides.

Diagnostic findings (recorded in `.scratch/diag-logs/`):
- baseline (SPEC=off):                tps=21, coherent
- nextn pre-fix (all NextN flags on): tps=4,  garbage
- step A (prime loop off):            tps=4,  garbage     (prime not the cause)
- step B (server overrides off):      tps=12, ~coherent   (main cause)

Safety net before next iteration; no fix applied yet.
Two server-context fixes required to make NextN speculative decoding
produce coherent target output and a small acceptance speedup on macOS
Metal. Pairs with cherry-pick 8ce2b9e (upstream Metal GDN
keep_intermediates=true), which is the actual root cause for the
garbage logits.

* NextN draft must NOT flip cparams.embeddings=true on the target
  context. Doing so reroutes the target graph to emit pooled/embedding
  outputs in place of vocab logits and breaks sampling for every
  generated token. NextN has its own pre-norm channel
  (llama_set_embeddings_pre_norm + llama_get_embeddings_pre_norm_ith);
  only Gemma 4 MTP needs the embeddings flag.

* Skip override_arch when --model-draft points at a standalone
  *_NEXTN_ONLY.gguf whose general.architecture is already
  qwen35_mtp / qwen35moe_mtp. Avoids the double-mmap of the target
  file when target and draft are different paths.

Also adds scripts/extract-qwen36-nextn-gguf.py to produce a self-
contained NextN draft GGUF from a combined *_MTP.gguf for the
separate-draft path.

Verified on Qwen3.6-27B-UD-Q4_K_XL_MTP + Q4_K_XL_NEXTN_ONLY draft on
Apple Silicon: 24.4 tok/s with acc=87.5% (DM=2) vs 20.85 tok/s
no-spec baseline.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ocessing

This update introduces an asynchronous worker thread to enhance the NextN speculative decoding process. The worker overlaps draft computation with server-side token processing, improving efficiency. Key changes include:

- Added a worker thread managed by a mutex and condition variable for handling draft requests.
- Implemented a pipeline mechanism that allows the system to return results from previous drafts while processing new requests.
- Introduced environment variable control for enabling/disabling the pipeline.

This enhancement aims to optimize performance and reduce latency in generating coherent outputs during NextN processing.
… GGUF)

Previously, NextN draft contexts loaded a second llama_model from the same
combined *_MTP.gguf with override_arch=qwen35*_mtp. On Apple Silicon (mmap=true)
each llama_model creates its own MTLBuffer covering the full file, so the 22 GB
Qwen3.6-35B-A3B target was mapped twice (~44 GB) and OOMed on M4 Max (38 GB
unified memory): kIOGPUCommandBufferCallbackErrorOutOfMemory.

The target model already loads the NextN-layer tensors into its own layer table
(see LLM_ARCH_QWEN35{,MOE} loaders, `layer.nextn.*` on `i >= n_layer -
nextn_predict_layers`). The draft context can reuse them directly:

  - Add cparams.nextn_draft + llama_context_params.nextn_draft (default false).
  - Add LLM_GRAPH_TYPE_NEXTN; llama_context routes decode/reserve through this
    gtype when cparams.nextn_draft=true.
  - In llama_model::build_graph, dispatch QWEN35 / QWEN35MOE + gtype=NEXTN to
    the existing llm_build_qwen35*_nextn builders (graphs unchanged otherwise;
    swap build_attn_inp_kv() → build_inp_mem_hybrid()->get_attn() because the
    target's memory is hybrid attn+GDN, not pure KV — pure-KV cast was UB).
  - llama_context ctor temporarily flips hparams.kv_only_nextn=true around
    create_memory() so the draft's KV cache only allocates cells for the
    NextN layer; the target context (constructed earlier) keeps its own KV
    layout via the per-memory hparams copy.
  - llama_context::graph_params hands a tweaked hparams_eff to the graph
    builder for draft contexts so has_kv routes correctly.
  - llm_graph_input_mem_hybrid::set_input: guard the recurrent-state s_copy
    backend buffer; NextN graphs never reference it, so the scheduler doesn't
    allocate one.
  - server-context.cpp: when target has NextN tensors and --model-draft points
    at the same file, set speculative.model_dft = model and
    cparams_dft.nextn_draft = true (no llama_model_load_from_file). The legacy
    standalone NEXTN_ONLY GGUF path is preserved for users shipping the draft
    head as a separate artifact.
  - common_speculative_are_compatible_nextn accepts model_tgt == model_dft for
    the shared-model path.
  - Public API: llama_model_has_nextn_layer / llama_model_n_nextn_predict_layers.

Benchmarks (Apple M4 Max, Metal, prompt ~50 tokens, --draft-max=2 --draft-min=1,
ctx=8192, median of 3 runs; full table in NEXTN.md §7):

  qwen-35B-A3B MoE f16-nextn      long=512:  83.63 tps (+20.7% vs f16-base 69.30)
  qwen-35B-A3B MoE turbo3-nextn   long=512:  78.41 tps (+26.5% vs turbo3-base 61.97)

35B-A3B MoE no longer OOMs (one MTL0_Mapped buffer = 21784 MiB instead of two).
27B dense remains draft-compute-bound (NextN layer = full transformer block,
t_draft ≈ 2.6× t_verify on dense); async pipeline can't fully overlap, so
NextN is paritetical on f16 and ~-12% on turbo3 (turbo3 dequant inside NextN
attention adds ~7% draft compute). Documented as known limitation in NEXTN.md §7.

Co-authored-by: Cursor <cursoragent@cursor.com>
…model

This commit introduces a new script, `sanity-27b-turbo3-base.sh`, which performs a cold-start sanity check for the 27B turbo3-base model. The script measures throughput against a historical baseline of ~18.4 TPS to identify potential thermal or code regressions. Key features include server initialization, health checks, and performance measurement over multiple runs with a predefined prompt. The script aims to ensure the model's operational integrity and performance consistency.
Brings in Gemma 4 + TurboQuant KV cache fixes:
- fix/turbo-rope-shift-gemma4 (PR #10)
- fix/iswa-get-can-shift-gemma4 (PR #9)
- fix/mtp-assistant-tensor-prefix (PR #7)
…ma 4 MTP

- Updated benchmark logs for Qwen 3.6 NextN, showing improved throughput of +24-36% on MoE targets and +5-7% on dense models.
- Revised performance notes in NEXTN.md to reflect shared-model draft path optimizations, eliminating the need for a second mmap.
- Enhanced README.md with new Qwen 3.6 NextN features and usage instructions, including shared model configurations.
- Adjusted MTP.md to include updated matrix benchmarks and observations for Gemma 4, highlighting performance gains and acceptance rates.
- Improved clarity in documentation regarding the integration of TurboQuant KV with NextN for optimal performance.
@Ooooze
Ooooze merged commit 514e600 into feature/turboquant-kv-cache May 12, 2026
27 of 57 checks passed
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
* oai moe

* compat with new checkpoint

* add attn sink impl

* add rope scaling yarn

* logits match with latest transformers code

* wip chat template

* rm trailing space

* use ggml_scale_bias

* rm redundant is_swa_all

* convert interleaved gate_up

* graph : fix activation function to match reference (AtomicBot-ai#7)

* vocab : handle o200k_harmony special tokens

* ggml : add attention sinks support (AtomicBot-ai#1)

* llama : add attn sinks

* ggml : add attn sinks

* cuda : add attn sinks

* vulkan : add support for sinks in softmax

remove unnecessary return

* ggml : add fused swiglu_oai op (AtomicBot-ai#11)

* ggml : add fused swiglu_oai op

* Update ggml/src/ggml-cpu/ops.cpp

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* update CUDA impl

* cont : metal impl

* add vulkan impl

* test-backend-ops : more test cases, clean up

* llama : remove unfused impl

* remove extra lines

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

---------

Co-authored-by: slaren <slarengh@gmail.com>

* repack mxfp4 upon conversion

* clean up a bit

* enable thinking

* add quick hack to render only some special tokens

* fix bf16 conversion

* remove vocab hack

* webui ok

* support chat parsing for gpt-oss

* fix webui

* direct mapping mxfp4, FINALLY

* force using mxfp4

* properly use lazy tensor

* ggml : add mxfp4

ggml : use e8m0 conversion instead of powf

Co-authored-by: Diego Devesa <slarengh@gmail.com>

change kvalues_mxfp4 table to match e2m1 (AtomicBot-ai#6)

metal : remove quantization for now (not used)

cuda : fix disabled CUDA graphs due to ffn moe bias

vulkan : add support for mxfp4

cont : add cm2 dequant

* ggml : add ggml_add_id (AtomicBot-ai#13)

* ggml : add ggml_add_id

* add cuda impl

* llama : add weight support check for add_id

* perf opt

* add vulkan impl

* rename cuda files

* add metal impl

* allow in-place ggml_add_id

* llama : keep biases on CPU with --cpu-moe

* llama : fix compile error

ggml-ci

* cuda : add fallback for __nv_cvt_e8m0_to_bf16raw

ggml-ci

* cleanup

ggml-ci

* sycl : fix supports_op for MXFP4

ggml-ci

* fix Unknown reasoning format

* ggml-cpu : fix AVX build

ggml-ci

* fix hip build

ggml-ci

* cuda : add mxfp4 dequantization support for cuBLAS

ggml-ci

* ggml-cpu : fix mxfp4 fallback definitions for some architectures

ggml-ci

* cuda : fix version required for __nv_cvt_e8m0_to_bf16raw

---------

Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: slaren <slarengh@gmail.com>
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
…rg#17764)

* Squashed commit of the following:

commit b3c6bf4b0450d8d452b934df27a0fb7cb53cd755
Author: Abhijit Ramesh <abhijitramesh2k@gmail.com>
Date:   Mon Dec 1 18:29:00 2025 -0800

    ggml webgpu: fix xielu parameter passing (AtomicBot-ai#11)

    The XIELU operation was incorrectly using static_cast to convert
    float parameters to uint32_t, which converted numeric values instead
    of preserving IEEE 754 bit patterns. This caused incorrect values
    to be interpreted by the GPU shader.

    * Use reinterpret_cast to preserve float bit patterns when passing
      through uint32_t params buffer
    * Update WGSL shader parameter types from u32 to f32
    * Re-enable XIELU support (was disabled due to numerical issues)

    Fixes NMSE test failures for XIELU operation on WebGPU backend.

commit 5ca9b5e
Author: neha-ha <137219201+neha-ha@users.noreply.github.com>
Date:   Tue Nov 18 12:17:00 2025 -0800

    Refactored pipelines and workgroup calculations (AtomicBot-ai#10)

    * refactored pipelines

    * refactored workgroup calculation

    * removed commented out block of prior maps

    * Clean up ceiling division pattern

    ---------

    Co-authored-by: Neha Abbas <nehaabbas@eduroam-169-233-141-223.ucsc.edu>
    Co-authored-by: Reese Levine <reeselevine1@gmail.com>

Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 29 23:13:06 2025 -0700

    formatted embed wgsl and ggml-webgpu.cpp

commit e1f6bae
Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 29 23:08:37 2025 -0700

    implemented REPL_Template support and removed bug in unary operators kernel

commit 8c70b8f
Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 15 16:14:20 2025 -0700

    responded and dealt with PR comments

commit f9282c6
Author: James Contini <jamescontini@gmail.com>
Date:   Sun Oct 12 13:41:41 2025 -0700

    removed unnecesarry checking if node->src[1] exists for unary operators

commit 4cf28d7
Author: James Contini <jamescontini@gmail.com>
Date:   Sun Oct 12 13:32:45 2025 -0700

    All operators (inlcluding xielu) working

commit 74c6add
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 13:16:48 2025 -0700

    fixed autoconfig

commit 3627499
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 13:10:46 2025 -0700

    removed vestigial files

commit cb08583
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 12:59:32 2025 -0700

    abides by editor-config

commit 5360e28
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 12:45:57 2025 -0700

    rms_norm double declaration bug atoned

commit 7b09baa
Merge: 8a6ec84 1579702
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 11:50:03 2025 -0700

    resolving merge conflicts

commit 8a6ec84
Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 8 18:06:47 2025 -0700

    unary operators pass ggml tests

commit c3ae382
Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 1 16:22:40 2025 -0700

    neg passes backend test

commit aa1c9b2
Author: James Contini <jamescontini@gmail.com>
Date:   Tue Sep 30 23:55:27 2025 -0700

    neg f16xf32xip builds and runs, havent actually ran a model that uses neg kernel yet though

Co-authored-by: James Contini <jamescontini@gmail.com>
Co-authored-by: Neha Abbas <neabbas@ucsc.edu>
Co-authored-by: Abhijit Ramesh <abhijitramesh2k@gmail.com>

* Remove extra code and format

* Add ops documentation (finally)

* Update ggml/src/ggml-webgpu/wgsl-shaders/embed_wgsl.py

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

---------

Co-authored-by: James Contini <jamescontini@gmail.com>
Co-authored-by: Neha Abbas <neabbas@ucsc.edu>
Co-authored-by: Abhijit Ramesh <abhijitramesh2k@gmail.com>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
…d per-thread state (ggml-org#18976)

* Squashed commit of the following:

commit b3c6bf4b0450d8d452b934df27a0fb7cb53cd755
Author: Abhijit Ramesh <abhijitramesh2k@gmail.com>
Date:   Mon Dec 1 18:29:00 2025 -0800

    ggml webgpu: fix xielu parameter passing (AtomicBot-ai#11)

    The XIELU operation was incorrectly using static_cast to convert
    float parameters to uint32_t, which converted numeric values instead
    of preserving IEEE 754 bit patterns. This caused incorrect values
    to be interpreted by the GPU shader.

    * Use reinterpret_cast to preserve float bit patterns when passing
      through uint32_t params buffer
    * Update WGSL shader parameter types from u32 to f32
    * Re-enable XIELU support (was disabled due to numerical issues)

    Fixes NMSE test failures for XIELU operation on WebGPU backend.

commit 5ca9b5e
Author: neha-ha <137219201+neha-ha@users.noreply.github.com>
Date:   Tue Nov 18 12:17:00 2025 -0800

    Refactored pipelines and workgroup calculations (AtomicBot-ai#10)

    * refactored pipelines

    * refactored workgroup calculation

    * removed commented out block of prior maps

    * Clean up ceiling division pattern

    ---------

    Co-authored-by: Neha Abbas <nehaabbas@eduroam-169-233-141-223.ucsc.edu>
    Co-authored-by: Reese Levine <reeselevine1@gmail.com>

Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 29 23:13:06 2025 -0700

    formatted embed wgsl and ggml-webgpu.cpp

commit e1f6bae
Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 29 23:08:37 2025 -0700

    implemented REPL_Template support and removed bug in unary operators kernel

commit 8c70b8f
Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 15 16:14:20 2025 -0700

    responded and dealt with PR comments

commit f9282c6
Author: James Contini <jamescontini@gmail.com>
Date:   Sun Oct 12 13:41:41 2025 -0700

    removed unnecesarry checking if node->src[1] exists for unary operators

commit 4cf28d7
Author: James Contini <jamescontini@gmail.com>
Date:   Sun Oct 12 13:32:45 2025 -0700

    All operators (inlcluding xielu) working

commit 74c6add
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 13:16:48 2025 -0700

    fixed autoconfig

commit 3627499
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 13:10:46 2025 -0700

    removed vestigial files

commit cb08583
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 12:59:32 2025 -0700

    abides by editor-config

commit 5360e28
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 12:45:57 2025 -0700

    rms_norm double declaration bug atoned

commit 7b09baa
Merge: 8a6ec84 1579702
Author: James Contini <jamescontini@gmail.com>
Date:   Fri Oct 10 11:50:03 2025 -0700

    resolving merge conflicts

commit 8a6ec84
Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 8 18:06:47 2025 -0700

    unary operators pass ggml tests

commit c3ae382
Author: James Contini <jamescontini@gmail.com>
Date:   Wed Oct 1 16:22:40 2025 -0700

    neg passes backend test

commit aa1c9b2
Author: James Contini <jamescontini@gmail.com>
Date:   Tue Sep 30 23:55:27 2025 -0700

    neg f16xf32xip builds and runs, havent actually ran a model that uses neg kernel yet though

Co-authored-by: James Contini <jamescontini@gmail.com>
Co-authored-by: Neha Abbas <neabbas@ucsc.edu>
Co-authored-by: Abhijit Ramesh <abhijitramesh2k@gmail.com>

* Remove extra code and format

* Add ops documentation (finally)

* ggml webgpu: add SOFTPLUS unary operator

Implements SOFTPLUS (log(1 + exp(x))) with f16/f32 support. Uses f32
precision for intermediate calculations to prevent f16 overflow.

* Add shader implementation and 4 variants (f32/f16, inplace/non-inplace)
* Register pipelines and device support
* Follow Vulkan backend numerical stability pattern

* ggml webgpu: add EXPM1 unary operator

Implements EXPM1 (exp(x) - 1) with f16/f32 support.

* Add shader implementation and 4 variants (f32/f16, inplace/non-inplace)
* Register pipelines and device support

* ggml webgpu: add FLOOR unary operator

Implements FLOOR (rounds down to nearest integer) with f16/f32 support.

* Add shader implementation and 4 variants (f32/f16, inplace/non-inplace)
* Register pipelines and device support

* ggml webgpu: add CEIL unary operator

Implements CEIL (rounds up to nearest integer) with f16/f32 support.

* Add shader implementation and 4 variants (f32/f16, inplace/non-inplace)
* Register pipelines and device support

* ggml webgpu: add ROUND unary operator

Implements ROUND (rounds to nearest integer) with f16/f32 support.

* Add shader implementation and 4 variants (f32/f16, inplace/non-inplace)
* Register pipelines and device support

* ggml webgpu: add TRUNC unary operator

Implements TRUNC (truncates towards zero) with f16/f32 support.

* Add shader implementation and 4 variants (f32/f16, inplace/non-inplace)
* Register pipelines and device support

* docs : update WebGPU support for unary operators (FLOOR, CEIL, ROUND, TRUNC, EXPM1, SOFTPLUS)

* Updates to webgpu get_memory

* Move shared state (webgpu_context) and device creation out of registration context, device context, and buffer context, and move into backend context

* Small cleanup

* Move Instance, Device, Adapter, Device creation, and capabilities to global state while moving Queue, pipelines, and buffers to per-thread state.

* Cleanups

* More cleanup

* Move staging_buf mutex to global context

* Resolve merge

* Resolve merge

* Resolve merge

* Clean up merge errors, delete forward declaration, and run clang-format

* Rename device_init to backend_init

* Move webgpu_context to backend_context

* Move buffer context members into global context and refactor function calls

* Run clang-format

* Remove commends

* Move parameter buffers to per-thread, add single memset_tensor param buf

* Fix CI compilation issue

* Fix builds for emscripten not supporting subgroups

* cleanup

* cleanup

---------

Co-authored-by: Reese Levine <reeselevine1@gmail.com>
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
)

* ggml: backend-agnostic tensor parallelism

* support for GPT-OSS, Qwen 3 MoE

* partial Vulkan fix

* add support for 4/8 GPUs

* unconditional peer access

* re-use buffers + ggml contexts

* fix output pattern

* NCCL support

* GGML: HIP: add RCCL support

* Remove shfl and AllReduce from backend interface

* move allocation workaround out of ggml-alloc.c

* 2d tensor set/get support

* Fix the seg fault without NCCL

* Apply suggestion from JohannesGaessler

* support for tensor dims % n_devs != 0

* fix view_offs scaling

* arbitrary num. of GPUs/tensor split

* fix compilation

* better granularity estimate

* Support device-specific host buffer types if all underlying backends expose the same type. This allows using pinned memory instead of pageable memory for CUDA.

Fix compilation errors.

* partial Qwen 3 Next support

* Fix qwen3 30b (AtomicBot-ai#8)

* Fix crash with Qwen-30B-A3B Q4_0

Qwen-30B-A3B Q4_0 has an intermediate dimension of 768. Using a granularity of 256 forces an uneven split between GPUs, which is not supported by the current implementation.

* Decide block size based on tensor quantization type

* Fix crashes due to KV cache serialization (AtomicBot-ai#9)

KV cache serialization requires non-zero offsets on the tensor. Add support in the meta backend to set/get a tensor with a non-zero offset.

* metal : fix build (AtomicBot-ai#7)

* static memory allocations, fix usage count

* fix tensor granularity

* more even memory distribution

* use BF16 for allreduce

* rebase fixup

* better error message for unsupported architectures

* Fix device mismatch during scatter of allReduce. (AtomicBot-ai#11)

There is a mismatch between the dst buffer device and the backend device, causing the use of sync copies

* Enable the previous allreduce implementation. It is better in both perf and stability (AtomicBot-ai#12)

* delay AllReduce for Moe for less I/O

* build : clean-up compile warnings

* backend : move most of the meta backend API to ggml-backend-impl.h

* cont : hide unused public API in the implementation

* llama : use llama_device + remove ggml_backend_dev_is_meta()

* ggml-backend : remove unused alloc include

* minor : remove regex include

* ggml : introduce ggml-ext.h for staging new APIs

* rebase fixup

* fix tests

* llama : more robust logic for determining Meta devices (AtomicBot-ai#16)

* llama : more robust logic for determining Meta devices

* cont : fix devs size check

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* cont : fix log type

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* disable roundtrip for meta backend

* fix arch selection

* Qwen 3.5 support

* fix Gemma 4 MoE

* fix OpenVino, SYCL

* fix test-llama-archs for CPU-only builds

* Fix Qwen 3.5 MoE

* disable meta backend tests for WebGPU

* tests : filter CPU-based devices from the Meta backend tests (AtomicBot-ai#17)

* meta : formatting, naming, indentation (AtomicBot-ai#18)

* formatting : llama-model.cpp

* formatting : ggml-ext.h

* formatting : ggml-backend-meta.cpp

* meta : add TODO

* add documentation

* better error messages

* fix GPT-OSS

---------

Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz>
Co-authored-by: Gaurav Garg <gaugarg@nvidia.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
Complete experiment log:
  AtomicBot-ai#1  4-mag LUT:           15.1 at 8K (BEST, +38%)
  AtomicBot-ai#2  Batched extract:     13.7 (+25%)
  AtomicBot-ai#3  Inline FA block:     13.5 (I-cache pressure)
  AtomicBot-ai#4  Deferred norm:       12.9 (loses ILP)
  AtomicBot-ai#5  2-pair half2:        12.0 (ternary overhead)
  AtomicBot-ai#6  Select chain:        11.9 (branches kill)
  AtomicBot-ai#7  Bit-arithmetic:      11.6 (ALU too heavy)
  AtomicBot-ai#8  FMA branchless:      11.4 (ALU still too heavy)
  AtomicBot-ai#9  Named-reg ternary:   10.3 (branches worst)
  AtomicBot-ai#10 Main (8-LUT):        10.95 (baseline)
  AtomicBot-ai#11 Non-vec FA:          10.2 (wrong kernel)
  Ceiling:                 24.5 (no dequant)

Apple8 hardware truth:
  1 divergent constant read < 7 ALU ops (even with fma)
  Branches cost MORE than divergent constant reads
  Array indexing ALWAYS spills on Metal
  4 constant addresses is the sweet spot

The 4-mag LUT is the dequant-level ceiling on Apple Silicon.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: tturney@psyguard.ai
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 12, 2026
Systematische Prüfung aller ROADMAP-Items gegen Fork-Code:
- AtomicBot-ai#11 EAGLE-3 (PR ggml-org#18039): bereits integriert (Commit 5777425)
- AtomicBot-ai#12 Coopmat2 (PR ggml-org#19075): bereits integriert (flash_attn_cm2.comp, SPV generiert)
- AtomicBot-ai#20 Tensor Parallelism (PR ggml-org#19378): bereits integriert (Commit d850df3)
- AtomicBot-ai#28 Adaptive MTP (PR ggml-org#22931): Fork hat eigene Implementierung
  (LLAMA_MTP_SKIP_STREAK_THRESHOLD), PR closed/inkompatibel

Meilenstein-Status:
- M1: ✅ abgeschlossen
- M2: ✅ evaluiert (AtomicBot-ai#6✅, AtomicBot-ai#7✅, AtomicBot-ai#12✅ bereits integriert, AtomicBot-ai#9❌, AtomicBot-ai#10❌)
- M3: ⏳ teilweise (AtomicBot-ai#3✅, AtomicBot-ai#13❌, AtomicBot-ai#14⏭️, AtomicBot-ai#15 offen)
- M4: ✅ abgeschlossen (AtomicBot-ai#11✅, AtomicBot-ai#28✅ eigene Implementierung)
- M5: ⏳ teilweise (AtomicBot-ai#12✅, AtomicBot-ai#20✅ bereits integriert, AtomicBot-ai#21 offen)
- M6: ☐ offen (Tier 4 Forschung)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants