Skip to content

serve: one slot for Qwen3.5/3.8 on HRX when --parallel is not given - #244

Merged
bong-water-water-bong merged 1 commit into
mainfrom
hrx-gdn-one-slot
Sep 30, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
hrx-gdn-one-slot

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Fixes HTTP 500s on 1bit serve --device hrx for Qwen3.5/3.8 models (qwen35, qwen35moe, qwen3next, including Ternary Bonsai) when two requests overlap and --parallel isn't given.

The failure: without --parallel, serve passes no -np, so llama-server opens its default several slots. A second concurrent request then shares a batch with the first. That hits the multi-sequence softplus / GATED_DELTA_NET path, which has no HRX kernel yet, and fails with HTTP 500. The leaderboard session found it: its Bonsai GSM8K run died on its second request.

The fix: for those architectures on HRX with no --parallel, serve now passes -np 1. Requests queue instead of failing.

Related PRs:

Verification: serve.cpp passes a syntax-only compile against the vendored headers. I didn't run it end-to-end; CI builds it.

🤖 Generated with Claude Code

Without --parallel, serve passes no -np and llama-server opens its default
several slots. On HRX0 a second concurrent request then shares a batch with
the first, and Qwen3.5/3.8 (qwen35, qwen35moe, qwen3next; Ternary Bonsai too)
fail with HTTP 500 at the multi-sequence softplus / GATED_DELTA_NET (found by
the leaderboard session: Bonsai GSM8K died on its second request). serve now
passes -np 1 for those architectures on HRX; requests queue instead.
Complements #243 (-np 1 for an explicit --parallel 1) and #240 (--parallel >1
refused for them on HRX).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit 65d2944 into main Sep 30, 2026
8 of 9 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-gdn-one-slot branch September 30, 2026 04:08
bong-water-water-bong added a commit that referenced this pull request Sep 30, 2026
… route, pin llama.cpp b8d587e (#248)

* Ternary Bonsai on HRX: exact PTQ1_0 -> Q4_0 converter and the prism.hadamard route

PrismML's Ternary-Bonsai-2-27B (Qwen3.8-27B trained ternary) now runs on HRX0.

- tools/ternary_to_q4_0.py writes PrismML's PTQ1_0 weights as Q4_0, bit for bit
  (q = trit + 8, the group's fp16 scale in each of its four blocks), decodes every
  group back before keeping the file, keeps all metadata and stamps
  onebit.ternary_q4_0 = 128. tests/ternary_to_q4_0_test.py checks it against
  PrismML's own encoder (run in CI next to the Q4NX repack test).
- 1bit serve sends a prism.hadamard file to --device hrx (auto picks it; other
  devices are refused, upstream llama.cpp would ignore the rotation and answer
  garbage) and refuses PrismML's own ternary types with the converter's name.
  tests/prism_route.sh.
- docs/hrx.md: how to run it, what was checked, and the one-sequence limit of the
  Qwen3.5 / Qwen3.8 gated delta net on HRX0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp b8d587e: load Hadamard-folded weights (llama.cpp #34)

acf9c74 -> b8d587e is llama.cpp #34 only: src/llama-hadamard.{h,cpp} reads
prism.hadamard.* and rotates the activations before each folded weight, so
Ternary Bonsai runs on HRX0. docs/hrx.md: the Qwen3.5 / Qwen3.8 one-sequence
note now points at serve's one-slot default and --parallel refusal (#244).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* registry: regenerate for the llama.cpp b8d587e pin

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant