Skip to content

release: windows-x64-rocm archive, AMD ROCm bench plan (b10269-1.6.0) - #79

Merged
Vect0rM merged 10 commits into
masterfrom
dev
Sep 10, 2026
Merged

Vect0rM merged 10 commits into
masterfrom
dev

Conversation

@Vect0rM

@Vect0rM Vect0rM commented Sep 10, 2026

Copy link
Copy Markdown
Member

Promote dev to master for b10269-1.6.0. What lands:

  • Windows AMD ROCm archive llama-turboquant-windows-x64-rocm.zip (self-contained, signed, beta until run on a Radeon)
  • Linux ROCm archive adds gfx950 / gfx1150 / gfx1103, dead rocWMMA flag dropped
  • docs/amd bench plan and scripts for ROCm vs Vulkan with TurboQuant
  • CHANGELOG section for b10269-1.6.0 (after ci(windows-rocm): sign the whole HIP payload + 1.6.0 changelog #78)

Same commits already built on dev; dev-latest 2026-09-09 carries the new archive.

🤖 Generated with Claude Code

Vect0rM and others added 9 commits September 8, 2026 15:00
New windows-x64-rocm job in dev-build and release: HIP backend built with
the clang from the ROCm 10.0 wheels (TheRock), consumer gfx targets only
(RDNA2-RDNA4, Ryzen AI 300, Strix Halo), --offload-compress. The publish
job glues the CPU archive with ggml-hip.dll, the HIP runtime and
rocBLAS/hipBLAS plus gfx-filtered Tensile kernels, so the zip loads with
only an Adrenalin driver (upstream stopped bundling rocBLAS and its zip
needs a ROCm SDK, ggml-org#26996).

ROCm CI hygiene: drop the dead GGML_HIP_ROCWMMA_FATTN flag (removed
upstream), add gfx950/gfx1150/gfx1103 to the Linux targets, fix the
stale CDNA comment, update the archive tables.

docs/amd: status of HIP in the fork, ROCm vs Vulkan bench matrix for
TurboQuant, roadmap. scripts/bench-amd.sh + bench-amd-report.py run the
matrix per backend build and render one table.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…iagnostics

Temporary workflow_dispatch-only copy of the windows-x64-rocm job: prints
where the ROCm 10 wheels keep the rocBLAS kernels, checks that
--offload-compress reaches clang, silences -Wignored-attributes, and makes
the runtime bundling non-fatal. To be deleted before merge.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…l stores

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The ROCm 10 Windows wheels ship no rocBLAS kernel library; rocblas.dll
imports libhipblaslt.dll and kernels come from amd_comgr at run time.
Walk the DLL imports with dumpbin instead of looking for a library dir.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- --offload-compress goes through CMAKE_CXX_FLAGS: on Windows ggml-hip
  compiles the .cu files as CXX via hip::device, so CMAKE_HIP_FLAGS is
  ignored. ggml-hip.dll drops from 545 MB to 47 MiB for 9 targets.
- Bundle the import closure of ggml-hip.dll (amdhip64_7, rocm_kpack,
  amd_comgr, hipblas, rocblas, libhipblaslt, rocsolver, origami; 230 MiB,
  92 MiB zipped) instead of a rocBLAS kernel library: the ROCm 10 wheels
  ship none, kernels come from amd_comgr at run time.
- -Wno-ignored-attributes: 850k GGML_API dllimport warnings made the job
  log unfetchable.
- Drop the scratch workflow used for iteration.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
ci: windows-x64-rocm archive + AMD ROCm bench plan
Sign hip-out after bundling, the same rule as the CUDA archives: the
code-sign action skips DLLs that already carry a valid signature, so
AMD-signed DLLs stay AMD's and unsigned ones get ours.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation devops labels Sep 10, 2026
ci(windows-rocm): sign the whole HIP payload + 1.6.0 changelog
@Vect0rM
Vect0rM merged commit 97211a0 into master Sep 10, 2026
48 of 55 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

devops documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant