Skip to content

server: preserve RAM prompt cache across slot reuse - #2

Merged
novkien merged 9 commits into
masterfrom
design/issue-1-kv-cache-retention
Sep 30, 2026
Merged

novkien merged 9 commits into
masterfrom
design/issue-1-kv-cache-retention

Conversation

@novkien

@novkien novkien commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Preserve a conversation's reusable native state in RAM before another prompt destructively borrows its slot. Saved snapshots remain available after prefix borrowing and can restore into a different execution slot. Selection uses checkpoint-compatible reuse; cache opt-out and existing RAM budgets remain enforced.

Addresses #1. Related incident: novkien/llama-proxy#416.

The runtime repair retains the owner's original change and adds non-consuming snapshots, shared immutable checkpoint backing, safe capture/restore, compatibility checks and focused regressions. The current test follow-up replaces the obsolete idle-cache log assertion with a near-full restored-token assertion. Optional SSD spill and --cache-spill-mib are not implemented.

Validation: 14 focused HTTP retention/idle-cache cases and 33 native tiny-model planner/capture/restore checks pass. The prior repair's full review and CPU validation are recorded at #2 (comment). Native Qwen/Nex hybrid, actual draft, long-context and production acceptance are being completed under the owner's current deployment request; they are not inferred from this build or merge.

This is an owner-fork delivery with ChatGPT assistance. No upstream submission is included. Issue closure follows recorded runtime acceptance.

Addresses #1

Documentation-only workspace for shared automatic/pinned cache handling,
branch-preserving RAM snapshots and optional SSD spill. No runtime changes.

Assisted-by: ChatGPT
Save live branch when incoming tokens would discard it, instead of f_keep < 0.5 gate. Always try RAM lookup for completion tasks including pinned slots. Rank cache entries by longest reusable prefix, keep 0.25 keep guard.

novkien commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Review intake: runtime implementation is not present in this PR

Inspected head: 18a6e1e53c181d02dffa81d595e1c6795c505b12, branch design/issue-1-kv-cache-retention, against master at 526c43b8f7dfea9032e9f35e7a1be9183ca7cc20.

The owner requested a code review and direct repairs, stopping before the owner merges, builds and runs the candidate. I checked the current PR metadata, changed-file list, actual patch, discussion, repository PR listing and exact remote head ref.

This PR still contains only the original five added files under docs/kv-cache-retention/: README.md, DESIGN.md, IMPLEMENTATION.md, TESTING.md and RESEARCH.md (one commit, 609 additions). There are no native source or executable regression-test changes to review or repair at this head. No different implementation PR appeared in the repository PR listing retrieved for this review; this does not establish the contents of unpublished local work.

GitHub currently reports mergeable=true, which establishes absence of Git merge conflicts, not implementation completeness or runtime correctness. Merging and building this head will not apply the KV-cache retention repair described in issue #1.

No source change, native build/inference test, ready-for-review transition or merge was performed. Keep the PR draft for the intended implementation delivery. The missing input is the implementation commit(s): publish them to this PR's head branch, or identify the implementation PR/branch so the review and repairs target the actual candidate.

novkien commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Code review resumed at 8cfa7bb

The owner has now published implementation commit 8cfa7bb52fe69e904abc53c86b775c01a6ec3b0b and explicitly requests direct review repairs on this PR, stopping before their merge/build/live test. The prior docs-only intake is superseded for this head.

The new changes affect server-context.cpp and server-task.cpp. I am checking preservation, non-consuming restore, checkpoint-aware ranking, memory/exception safety, cache opt-out and busy/parent admission, with focused regressions. Initial inspection confirms load still consumes the RAM source, and ranking still uses raw LCP with the 25% exclusion rather than valid checkpoint reuse. No SSD implementation is present at this head; the SSD design is not being represented as working code.

Review and repair scope: this PR's native RAM-cache candidate and its necessary tests/documentation. No merge, production restart, proxy-policy change or upstream submission. Validation will identify the exact tested head and any unrun native/hardware checks.

@github-actions github-actions Bot added documentation Improvements or additions to documentation devops server labels Sep 29, 2026
novkien and others added 5 commits September 29, 2026 10:40
The reviewed RAM source was built, tested and published as e8ef91d. No permanent workflow change is needed.

Addresses #1
Assisted-by: ChatGPT

novkien commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Review repairs published: RAM candidate ready for owner build/live trial

Reviewed owner input: 8cfa7bb52fe69e904abc53c86b775c01a6ec3b0b.
Repair commit: e8ef91dc691f1197f9a8f69b586563de377d6278.
Current delivery head: 705ee67a9aaf75409cad53c5a5c334ec61b64be9. The last commit only removes the temporary review transport workflow; comparison confirms no source/test change after the tested repair. Existing owner commits were retained; publication used a non-force push.

Findings repaired

  1. Borrowing a RAM prefix destroyed its source. load() no longer clears serialized bytes, moves the saved prompt away or erases a snapshot. It prepares independent live metadata and updates recency only after successful restore.
  2. Raw-LCP ranking and the 25% exclusion selected/missed the wrong state. server_prompt::plan_reuse extracts the existing checkpoint/position rule and is shared by candidate ranking and execution. Equal benefit favors live state; adapter mismatch is excluded. This does not invent missing recurrent checkpoints.
  3. Selection had side effects before admission was complete. Slot selection is now read-only, including explicit IDs with similarity disabled. Cache preparation runs in launch after availability/family-capacity and input checks. An old branch is captured before replacement under its old adapter identity.
  4. Cache opt-out was not respected. cache_prompt=false disables incoming reuse and automatic archiving of that request. Infill keeps its existing live reuse behavior without becoming an automatic RAM archive.
  5. A failed optional idle save still cleared the only live state. Optional evacuation retains live state when preservation fails. A required KV-pressure purge attempts preservation and explicitly logs unavoidable loss.
  6. Checkpoint copies and incomplete capture could exhaust/corrupt the cache. Large target/draft checkpoint buffers now have immutable shared backing; a new capture allocates fresh storage. Full entry allocation is exception-guarded and serialized byte counts must match before publication. Exact compatible snapshots are deduplicated; token-prefix containment alone no longer deletes another branch. A selected restore source is protected during capture/eviction.
  7. Restore metadata/failure handling needed tightening. Metadata is prepared before native setters; missing draft state is rejected before installation; source bytes survive failure. The caller clears only the failed destination and returns an explicit error rather than generating from mixed state.

Executed checks

  • New maintained HTTP cache fixtures: 12 PASS on repaired bytes; 7 FAILED / 5 PASS on the owner's initial native bytes. The seven baseline failures are borrowed-snapshot loss (automatic/pinned x similarity 0/0.1), the 25% exclusion, opt-out archiving, and failed idle save followed by clear.
  • Expanded existing completion plus manual text-slot suite: 53 PASS, 1 skipped, 1 slow model-specific test excluded. The tiny model was not substituted as evidence for the excluded Phi case.
  • New C++ test: 18 model-independent checks PASS; 33 checks total PASS with the public tiny model, including real native save/restore, cross-slot reuse, retained source, exact duplicate, protected byte-budget rejection and missing-draft rejection.
  • GCC 14 CPU Release build with fatal warnings: PASS. Additional test-recurrent-state-rollback, test-save-load-state, test-chat targets build; CTest test-chat and test-server-prompt-cache pass.
  • git diff --check and exact 13-file before/after blob manifest: PASS.
  • A fresh GitHub runner independently built the exact repair, ran CTest, 33 native checks and all 12 HTTP cache cases, then pushed that tested commit. Validation/publication run. Its cache-review-validation artifact records candidate SHA = published SHA = e8ef91dc691f1197f9a8f69b586563de377d6278 and the complete file manifest.

Model: stories260K.gguf, SHA-256 270cba1bd5109f42d03350f60406024560464db173c0e387d91f0426d3bd256d. Baseline comparison used the baseline's own shared libraries. The local harness timing overrides are not production performance evidence.

Scope and remaining evidence

This is a RAM-cache repair, not the complete RAM+SSD design. Optional SSD spill, persistent snapshot codec and --cache-spill-mib remain unimplemented; existing manual slot-file semantics are unchanged. Real Nex/hybrid GPU, actual draft/MTP/media, exhaustive fault injection and matched 100k-context performance remain PARTIAL until tested on their applicable configurations. CPU planner fixtures are not GPU acceptance.

README/DESIGN/IMPLEMENTATION/TESTING now distinguish delivered RAM behavior from the unfinished full architecture. Issue #1 must remain open. No merge, production restart, proxy pinning activation or upstream submission was performed.

The inherited CANN workflow has only commented-out jobs and has emitted invalid/failed workflow runs; it was not modified or misreported as a passing native test. Other inherited platform workflows are separate from the completed focused CPU evidence above. Git mergeability and all-platform CI success are different claims.

Owner endpoint: merge this reviewed RAM candidate when satisfied, build using the existing native build procedure, then run the bounded A/B/A and pinned-after-evacuation live checks. Retain current RAM/context limits and do not enable the proposed SSD flag, which is not implemented.

@novkien novkien changed the title docs: KV cache retention design and junior implementation workspace server: preserve RAM prompt cache across slot reuse Sep 29, 2026
@novkien
novkien marked this pull request as ready for review September 29, 2026 04:24
@novkien

novkien commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Pre-merge handoff: RAM repair and current owner trial

Addresses #1; linked incident novkien/llama-proxy#416.

The owner directly authorized deployment/build of the fork on jarvis-llm and live testing of PR #2 together with selective-P2P issue #3 on 2026-09-30. The scope here is the implemented RAM repair. Optional SSD spill remains unimplemented and is not claimed by this delivery.

Rechecked head 705ee67a9aaf75409cad53c5a5c334ec61b64be9; its source tree equals the clean host's local merge 2138591d22f9e1e203630afcac2181abcd96d068. The Ubuntu/Windows server CI failures were both the existing idle-cache test's assertion for the removed updating prompt cache log line. The reviewed test-only follow-up now drains the log and requires restored cache_n >= original_prompt_n - 1, retaining the lower-prefill and neighboring idle-slot checks. It does not alter native runtime code.

Verification of the reviewed source/test candidate: the 12 maintained RAM-retention HTTP cases plus both idle-cache cases pass (14 passed); the native tiny-model capture/restore, immutable ownership, budget, cross-slot and rewind fixture passes 33 checks. The tiny model is ggml-org/test-model-stories260K/stories260K-f32.gguf, SHA256 270cba1bd5109f42d03350f60406024560464db173c0e387d91f0426d3bd256d. Tests ran in a disposable Python environment with local model input and disposable server ports. git diff --check passes.

Merge authorizes source delivery; it is not cache acceptance on the production model. After the reviewed merge, the main task will build a paired native/RPC release with issue #3, verify actual Qwen/Nex hybrid and draft behavior, then record applicable runtime acceptance. Keep issue #1 open until that evidence is available. No SSD, upstream submission or unrelated source/configuration change belongs to this PR.

Reviewed published head: 12ac804.

@novkien
novkien merged commit 7865d7f into master Sep 30, 2026
15 of 31 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

devops documentation Improvements or additions to documentation server testing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant