Skip to content

Bump ZINC: zolotukhin/zinc b9123c649f81 - #39

Merged
bong-water-water-bong merged 1 commit into
mainfrom
bump-zinc/b9123c649f81
Sep 24, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
bump-zinc/b9123c649f81

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/zinc from c50b4add3402 to upstream main b9123c649f81.

Upstream commits (last 30):

  • ac84b5de5 rocm: hipBLAS prefill attention for Gemma 4
  • 34ebf0fd6 site: ROCm ahead of llama.cpp HIP in all 48 rows
  • 84b8ddf06 site: re-measure Vulkan Gemma 4 26B-A4B and Qwen 3.6
  • 6fc514024 vulkan: gather Gemma prefill embeddings on the GPU
  • bb5be230d vulkan: record several Gemma MoE prefill layers per submission
  • 6e6650b50 site: re-measure Vulkan Gemma 4 26B-A4B after the prefill changes
  • 46341e40d vulkan: overlap Gemma 4's shared expert with the routed experts in prefill
  • d407abaeb site: re-measure Vulkan Gemma 4 26B-A4B with the shared-expert prefill overlap
  • 2125e476c tests: widen Gemma grouped MoE prefill source windows
  • 25e9b81eb vulkan: record up to 16384 token-layers per Gemma MoE prefill submission
  • 1926f4564 vulkan: let the Gemma MoE prefill combine overwrite instead of accumulate
  • f447b5969 vulkan: batch the last layer of the Gemma MoE prefill
  • cffbd6dc3 site: re-measure Vulkan Gemma 4 26B-A4B after batching the last prefill layer
  • e1c1c9e05 tokenizer: merge Gemma 4 BPE by rank and always prepend BOS, like the reference
  • 60c753bc7 vulkan: fix Gemma MoE prefill gate/up dropping half of every Q4_K block
  • 852910afe bench: flag prompt token mismatches larger than one token
  • 9c9edc03f vulkan: run the Gemma MoE prefill gate/up on DP4a with Q8_1 activations
  • a2ee308e5 site: re-measure Gemma 4 on the R9700 after the tokenizer and gate/up fixes
  • bd7ee7323 vulkan: chain greedy decode steps so the GPU never waits for the host
  • 44d35047c vulkan: chain Muse Glimmer decode too
  • cd66b9c48 site: re-measure the Vulkan tab with chained greedy decode
  • d7a5ed9bf rocm: run Gemma's BLAS prefill attention in f16 on the matrix cores
  • b9123c649 site: re-measure ROCm Gemma 4 with f16 prefill attention

CI here has no GPU. Before merging, on Strix Halo:
scripts/build-zinc.sh ~/.cache/zinc-pin vulkan && RADV_PERFTEST=coop_matrix ~/.cache/zinc-pin/vulkan/bin/zinc -m ~/models/Qwen3-0.6B-Q4_K_M.gguf --prompt 'The capital of France is' (first token 12095).

@bong-water-water-bong
bong-water-water-bong merged commit fbd7b57 into main Sep 24, 2026
1 check passed
@bong-water-water-bong
bong-water-water-bong deleted the bump-zinc/b9123c649f81 branch September 24, 2026 12:22
bong-water-water-bong added a commit that referenced this pull request Oct 1, 2026
…260)

Wrapped lines in docs/hrx.md starting with '#36', '#39', '#40' became
<h1> headings (python-markdown needs no space after '#'), so hrx.html
had 4 h1s and the daily SEO check flagged it. Prefix them with 'PR'.
tools/site.py now fails the build when a page has other than one h1.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant