Repository navigation
NPU lane: fresh hw_context per generate(), fixes the idle timeout - #62
Conversation
…r times out After more than ~8 s idle the lane's long-lived context failed its next runlist with ERT_CMD_STATE_TIMEOUT and left a stuck context behind for 1-2 min. Lane::begin()/end() now create the context per generate() and drop it warm. lane_idle2: 10/10 rounds at 15 s and at 120 s idle, 97-101 tok/s, no contexts left. Root cause and fix by pi agent-74509b (goal mugi4zva). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍(Review updated until commit a594d00)Here are some key observations to aid the review process:
|
|
Persistent review updated to latest commit a594d00 |
…gram cache cap; serve runs PQ2_0/PTQ1_0 files on HRX (#280) llama.cpp fork since bd5b297: - #62: PrismML's PQ2_0 / PTQ1_0 as native ggml types (142 / 143) with HRX decode on the K-quant kernels; Q1_0 decode moves onto them. Ternary-Bonsai-2-27B (balanced mode, llama-bench -fa 1): PTQ1_0 5.53 GiB 14.3 tok/s, PQ2_0 6.70 GiB 15.8 tok/s, vs 14.13 GiB / 15.3 for the Q4_0 copy; KLD vs CPU 0.000129 for all three. CPU decode bit-identical to PrismML's build on every tensor of both 27B files. test-backend-ops -b HRX0 1020/1020. - #63: GGML_HRX_GRAPH_PROGRAM_CACHE (default 64) bounds HRX server memory: ZAYA1-8B over 40 varying-length requests peaks at 8.7 GiB instead of growing past 24.6 GiB; answers identical. serve: a file in PrismML's ternary types goes to HRX with --device auto, rotated (prism.hadamard) or not (Ternary-Bonsai-1.7B); another --device is refused with the converter's name. Smoke test on strixhalo: `1bit serve -m Ternary-Bonsai-2-27B-PTQ1_0.gguf` routed to HRX and answered "Paris" 3/3. tests/prism_route.sh covers the new routes; ctest 19/19 (build without HRX). Docs: docs/hrx.md Ternary Bonsai section (native types, table), the cache cap section, "Our patches". Registry regenerated (no mapping changes). Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The RAG servers started a Vulkan llama-server whatever the build. With HRX (ONEBIT_HRX_SERVER) they now run on HRX0 with one slot and 2048-token inputs: HRX runs one sequence per batch (docs/hrx.md), and 2048 is the largest prompt chunk its matmul kernels take in one pass. A build without HRX keeps Vulkan, 4 slots, 8192 tokens. /v1/models reports the real device. Checked on Strix Halo against Vulkan0 (HRX build of fork d60cc4f + #62), two repeats each: - Qwen3-Embedding-0.6B Q8_0: HRX repeats are bit-identical; cosine to Vulkan 0.99961 / 0.99988 / 0.99987 on three inputs. - bge-reranker-v2-m3 Q8_0: HRX repeats identical; scores 8.609 / -6.756 / -0.401 / -11.020 vs Vulkan 8.614 / -6.757 / -0.361 / -11.019, same order. jina-reranker-v1-tiny still fails on HRX (docs/serve.md). ctest 19/19. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
After about 8 s idle, the NPU lane's single long-lived
hw_contextfailed its next runlist withERT_CMD_STATE_TIMEOUT. It also left a stuck context that blocked every NPU user for 1–2 minutes, which broke1bit serve's NPU route between requests.Fix:
Lane::begin()creates the context (initload_pdirun + lm-head config) at the start of eachgenerate(), andLane::end()drops it while warm, also on the exception path. Weights, KV and activation buffers are independent of the context and stay allocated once.Measured on strixhalo (
lane_idle2, Qwen3-0.6B, runtime PM at its defaultcontrol=auto):xrt-smishows no hardware contexts.Before the fix, round 1 onwards failed at 8 s and at 15 s idle.
Follow-ups:
step()directly, so it needsbegin()/end()when it lands.npu/lax.cpp) also keeps long-lived contexts and has not been tested for idle yet.Root cause and patch: pi agent-74509b (goal mugi4zva); docs note added in
docs/npu.md.🤖 Generated with Claude Code