Repository navigation
serve: --parallel 1 starts llama-server with one slot - #243
Merged
Merged
Conversation
--parallel 1 was dropped (only values above 1 reached llama-server), so the
child ran with llama-server's own default of several slots. On HRX0 a second
slot can then share a batch, and Qwen3.5 / Qwen3.8 (Bonsai too) have no
multi-sequence gated delta net there: the request fails with HTTP 500
("unsupported HRX node ... UNARY" in the child's log). Found running GSM8K on
Ternary-Bonsai-2-27B through `1bit serve --device hrx --parallel 1`.
tests/parallel_args.sh (no GPU): --parallel 1 -> -np 1, no --parallel -> no
-np, --parallel 4 -> -np 4.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
bong-water-water-bong
added a commit
that referenced
this pull request
Sep 30, 2026
…244) Without --parallel, serve passes no -np and llama-server opens its default several slots. On HRX0 a second concurrent request then shares a batch with the first, and Qwen3.5/3.8 (qwen35, qwen35moe, qwen3next; Ternary Bonsai too) fail with HTTP 500 at the multi-sequence softplus / GATED_DELTA_NET (found by the leaderboard session: Bonsai GSM8K died on its second request). serve now passes -np 1 for those architectures on HRX; requests queue instead. Complements #243 (-np 1 for an explicit --parallel 1) and #240 (--parallel >1 refused for them on HRX). Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
1bit serve --parallel 1never reached llama-server:-npwas passed only for values above 1, so the child ran with llama-server's own default of several slots.On HRX0 a second slot can then share a batch, and Qwen3.5 / Qwen3.8 (Ternary Bonsai too) have no multi-sequence gated delta net there, so the request fails with HTTP 500. The child's log says
unsupported HRX node ... UNARY. I found it running GSM8K on Ternary-Bonsai-2-27B through1bit serve --device hrx --parallel 1: the speed prompts passed, then the second request failed.Now
--parallel 1passes-np 1. Without--parallel, llama-server's default still applies. That default is still exposed to the same failure on HRX with these models; that belongs with the--parallelrefusal for qwen35/qwen3next on HRX.Test.
tests/parallel_args.sh(no GPU, fake backend) checks:--parallel 1passes-np 1;--parallelpasses no-np;--parallel 4passes-np 4.Run on strixhalo together with
hadamard_routeandlong_route: 3/3 pass.🤖 Generated with Claude Code