Skip to content

NPU lane: fresh hw_context per generate(), fixes the idle timeout - #62

Merged
bong-water-water-bong merged 2 commits into
mainfrom
npu-idle-fix
Sep 25, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
npu-idle-fix

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

After about 8 s idle, the NPU lane's single long-lived hw_context failed its next runlist with ERT_CMD_STATE_TIMEOUT. It also left a stuck context that blocked every NPU user for 1–2 minutes, which broke 1bit serve's NPU route between requests.

Fix: Lane::begin() creates the context (init load_pdi run + lm-head config) at the start of each generate(), and Lane::end() drops it while warm, also on the exception path. Weights, KV and activation buffers are independent of the context and stay allocated once.

Measured on strixhalo (lane_idle2, Qwen3-0.6B, runtime PM at its default control=auto):

  • 15 s idle: 10/10 rounds, 97–101 tok/s (re-verified on this branch).
  • 120 s idle: 10/10 rounds, 94–100 tok/s (pi agent-74509b's run).
  • Afterwards xrt-smi shows no hardware contexts.

Before the fix, round 1 onwards failed at 8 s and at 15 s idle.

Follow-ups:

  • The npu-draft bench drives the lane with step() directly, so it needs begin()/end() when it lands.
  • The 35B path (npu/lax.cpp) also keeps long-lived contexts and has not been tested for idle yet.

Root cause and patch: pi agent-74509b (goal mugi4zva); docs note added in docs/npu.md.

🤖 Generated with Claude Code

…r times out

After more than ~8 s idle the lane's long-lived context failed its next runlist
with ERT_CMD_STATE_TIMEOUT and left a stuck context behind for 1-2 min.
Lane::begin()/end() now create the context per generate() and drop it warm.
lane_idle2: 10/10 rounds at 15 s and at 120 s idle, 97-101 tok/s, no contexts left.
Root cause and fix by pi agent-74509b (goal mugi4zva).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit a594d00

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit a594d00)

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Context Management in generate()

The new generate() function now calls lane.begin() at the start and lane.end() at the end (including in exception paths). This ensures a fresh hardware context is used for each generation, preventing the timeout issue caused by idle contexts. However, this change assumes that generate_once does not rely on any state that might be invalidated by the context recreation. If generate_once or any of the functions it calls (like step) depend on prior context state, this could lead to incorrect behavior or errors.

GenerateResult generate(Lane& lane, const Model& model, const std::vector<int>& prompt, const GenerateOptions& opt) {
    lane.begin();
    try {
        GenerateResult r = generate_once(lane, model, prompt, opt);
        lane.end();
        return r;
    } catch (...) {
        lane.end();
        throw;
    }
}
Hardware Context Lifecycle

The begin() and end() methods in Lane::Impl now manage the lifetime of the hw_context. The begin() method creates a new context and configures it with the init and lm-head configurations, while end() resets runlists, kernels, and the context itself. This approach correctly addresses the idle timeout issue, but care must be taken to ensure that all resources managed by the context are properly cleaned up in end() to avoid resource leaks or dangling references.

void begin() {
    const Bytes init = assemble_full_elf(ctx1, pdi, "flinit", PdiMode::kInitOnly);
    ctx = std::make_unique<xrt::hw_context>(dev, xrt::elf(reinterpret_cast<const char*>(init.data()), init.size()));
    {
        xrt::ext::kernel k(*ctx, "flinit");
        xrt::ext::bo unused(dev, 4096);
        xrt::run r(k);
        r.set_arg(0, unused);
        r.start();
        if (r.wait(std::chrono::milliseconds(5000)) != ERT_CMD_STATE_COMPLETED)
            throw std::runtime_error("NPU init run did not complete");
    }
    ctx->add_config(xrt::elf(reinterpret_cast<const char*>(head.data()), head.size()));
    lmhead = std::make_unique<xrt::ext::kernel>(*ctx, "flhead");
    layer_kernels.clear();
    slots[0].rl.reset(); slots[0].runs.clear();
    slots[1].rl.reset(); slots[1].runs.clear();
}

// Destroy the hw_context while it is warm (every run completed), so its
// destructor never waits on a stuck command. Drop the runlists, runs,
// layer kernels and lm-head kernel before the context they reference.
void end() {
    slots[0].rl.reset(); slots[0].runs.clear();
    slots[1].rl.reset(); slots[1].runs.clear();
    layer_kernels.clear();
    lmhead.reset();
    ctx.reset();
}

@1bit-MONSTER 1bit-MONSTER deleted a comment from github-actions Bot Sep 25, 2026
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit a594d00

@bong-water-water-bong
bong-water-water-bong merged commit d2a67f7 into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the npu-idle-fix branch September 25, 2026 08:52
bong-water-water-bong added a commit that referenced this pull request Oct 2, 2026
…gram cache cap; serve runs PQ2_0/PTQ1_0 files on HRX (#280)

llama.cpp fork since bd5b297:
- #62: PrismML's PQ2_0 / PTQ1_0 as native ggml types (142 / 143) with HRX decode on the K-quant
  kernels; Q1_0 decode moves onto them. Ternary-Bonsai-2-27B (balanced mode, llama-bench -fa 1):
  PTQ1_0 5.53 GiB 14.3 tok/s, PQ2_0 6.70 GiB 15.8 tok/s, vs 14.13 GiB / 15.3 for the Q4_0 copy;
  KLD vs CPU 0.000129 for all three. CPU decode bit-identical to PrismML's build on every tensor of
  both 27B files. test-backend-ops -b HRX0 1020/1020.
- #63: GGML_HRX_GRAPH_PROGRAM_CACHE (default 64) bounds HRX server memory: ZAYA1-8B over 40
  varying-length requests peaks at 8.7 GiB instead of growing past 24.6 GiB; answers identical.

serve: a file in PrismML's ternary types goes to HRX with --device auto, rotated (prism.hadamard)
or not (Ternary-Bonsai-1.7B); another --device is refused with the converter's name. Smoke test on
strixhalo: `1bit serve -m Ternary-Bonsai-2-27B-PTQ1_0.gguf` routed to HRX and answered "Paris" 3/3.
tests/prism_route.sh covers the new routes; ctest 19/19 (build without HRX).

Docs: docs/hrx.md Ternary Bonsai section (native types, table), the cache cap section, "Our patches".
Registry regenerated (no mapping changes).

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Oct 2, 2026
The RAG servers started a Vulkan llama-server whatever the build. With HRX (ONEBIT_HRX_SERVER)
they now run on HRX0 with one slot and 2048-token inputs: HRX runs one sequence per batch
(docs/hrx.md), and 2048 is the largest prompt chunk its matmul kernels take in one pass. A build
without HRX keeps Vulkan, 4 slots, 8192 tokens. /v1/models reports the real device.

Checked on Strix Halo against Vulkan0 (HRX build of fork d60cc4f + #62), two repeats each:
- Qwen3-Embedding-0.6B Q8_0: HRX repeats are bit-identical; cosine to Vulkan 0.99961 / 0.99988 /
  0.99987 on three inputs.
- bge-reranker-v2-m3 Q8_0: HRX repeats identical; scores 8.609 / -6.756 / -0.401 / -11.020 vs
  Vulkan 8.614 / -6.757 / -0.361 / -11.019, same order.
jina-reranker-v1-tiny still fails on HRX (docs/serve.md). ctest 19/19.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant