Repository navigation
server: preserve RAM prompt cache across slot reuse - #2
Conversation
Addresses #1 Documentation-only workspace for shared automatic/pinned cache handling, branch-preserving RAM snapshots and optional SSD spill. No runtime changes. Assisted-by: ChatGPT
Save live branch when incoming tokens would discard it, instead of f_keep < 0.5 gate. Always try RAM lookup for completion tasks including pinned slots. Rank cache entries by longest reusable prefix, keep 0.25 keep guard.
Review intake: runtime implementation is not present in this PRInspected head: The owner requested a code review and direct repairs, stopping before the owner merges, builds and runs the candidate. I checked the current PR metadata, changed-file list, actual patch, discussion, repository PR listing and exact remote head ref. This PR still contains only the original five added files under GitHub currently reports No source change, native build/inference test, ready-for-review transition or merge was performed. Keep the PR draft for the intended implementation delivery. The missing input is the implementation commit(s): publish them to this PR's head branch, or identify the implementation PR/branch so the review and repairs target the actual candidate. |
Code review resumed at 8cfa7bbThe owner has now published implementation commit The new changes affect Review and repair scope: this PR's native RAM-cache candidate and its necessary tests/documentation. No merge, production restart, proxy-policy change or upstream submission. Validation will identify the exact tested head and any unrun native/hardware checks. |
Addresses #1 Assisted-by: ChatGPT
Review repairs published: RAM candidate ready for owner build/live trialReviewed owner input: Findings repaired
Executed checks
Model: Scope and remaining evidenceThis is a RAM-cache repair, not the complete RAM+SSD design. Optional SSD spill, persistent snapshot codec and README/DESIGN/IMPLEMENTATION/TESTING now distinguish delivered RAM behavior from the unfinished full architecture. Issue #1 must remain open. No merge, production restart, proxy pinning activation or upstream submission was performed. The inherited CANN workflow has only commented-out jobs and has emitted invalid/failed workflow runs; it was not modified or misreported as a passing native test. Other inherited platform workflows are separate from the completed focused CPU evidence above. Git mergeability and all-platform CI success are different claims. Owner endpoint: merge this reviewed RAM candidate when satisfied, build using the existing native build procedure, then run the bounded A/B/A and pinned-after-evacuation live checks. Retain current RAM/context limits and do not enable the proposed SSD flag, which is not implemented. |
Addresses #1
Pre-merge handoff: RAM repair and current owner trialAddresses #1; linked incident novkien/llama-proxy#416. The owner directly authorized deployment/build of the fork on jarvis-llm and live testing of PR #2 together with selective-P2P issue #3 on 2026-09-30. The scope here is the implemented RAM repair. Optional SSD spill remains unimplemented and is not claimed by this delivery. Rechecked head Verification of the reviewed source/test candidate: the 12 maintained RAM-retention HTTP cases plus both idle-cache cases pass (14 passed); the native tiny-model capture/restore, immutable ownership, budget, cross-slot and rewind fixture passes 33 checks. The tiny model is Merge authorizes source delivery; it is not cache acceptance on the production model. After the reviewed merge, the main task will build a paired native/RPC release with issue #3, verify actual Qwen/Nex hybrid and draft behavior, then record applicable runtime acceptance. Keep issue #1 open until that evidence is available. No SSD, upstream submission or unrelated source/configuration change belongs to this PR. Reviewed published head: 12ac804. |
Preserve a conversation's reusable native state in RAM before another prompt destructively borrows its slot. Saved snapshots remain available after prefix borrowing and can restore into a different execution slot. Selection uses checkpoint-compatible reuse; cache opt-out and existing RAM budgets remain enforced.
Addresses #1. Related incident: novkien/llama-proxy#416.
The runtime repair retains the owner's original change and adds non-consuming snapshots, shared immutable checkpoint backing, safe capture/restore, compatibility checks and focused regressions. The current test follow-up replaces the obsolete idle-cache log assertion with a near-full restored-token assertion. Optional SSD spill and
--cache-spill-mibare not implemented.Validation: 14 focused HTTP retention/idle-cache cases and 33 native tiny-model planner/capture/restore checks pass. The prior repair's full review and CPU validation are recorded at #2 (comment). Native Qwen/Nex hybrid, actual draft, long-context and production acceptance are being completed under the owner's current deployment request; they are not inferred from this build or merge.
This is an owner-fork delivery with ChatGPT assistance. No upstream submission is included. Issue closure follows recorded runtime acceptance.