Skip to content

cuda: retain routed experts in bounded cache - #739

Open
exochard wants to merge 3 commits into
antirez:mainfrom
exochard:perf/cuda-persistent-expert-cache
Open

cuda: retain routed experts in bounded cache#739
exochard wants to merge 3 commits into
antirez:mainfrom
exochard:perf/cuda-persistent-expert-cache

Conversation

@exochard

@exochard exochard commented Aug 7, 2026

Copy link
Copy Markdown

Depends on #737 and #738 (the safety and async-upload branches are required bases).

Context

This change is stacked on the SSD safety and async-upload changes. DeepSeek routing is data-dependent: the same experts can be selected again in later tokens, so a bounded persistent cache can avoid repeated SSD reads and PCIe transfers. This is separate from the compact one-layer upload buffer.

Change

  • Allocate one bounded persistent arena for complete routed-expert records.
  • Track logical (layer, expert) keys and physical slots with LRU replacement; do not retain an entry smaller than the deterministic 480-slot token working set.
  • Reserve space conservatively alongside the generic range cache and session working memory, with fallback when the arena or metadata cannot be initialized.
  • Load misses through the pinned upload path, then either dispatch directly from physical slots or copy into the existing compact table.
  • Remap selected IDs to physical slots for direct kernels and size sorted/tiled workspaces against the physical capacity.
  • Record compute-done events after routed kernels and fence slot reuse; if event setup fails, synchronize before reuse.
  • Preserve the compact streaming path when the persistent cache is unavailable or a load fails.

Validation

Builds completed with no compiler warnings:

make -B ds4 CUDA_HOME=/opt/cuda CUDA_ARCH=sm_86
make -B ds4 CUDA_HOME=/opt/cuda CUDA_ARCH=sm_89

git diff --check is clean. Target hardware is an RTX 4060 Ti 16 GiB (sm_89) and the workload is the DeepSeek V4 Flash 0731 IQ2XXS SSD-streamed GGUF. End-to-end isolated-branch GPU execution still needs to be run from a driver-visible shell; the user's integration measurements motivating this work were approximately 1.9 t/s generation on compact streaming and 4.96--5.81 t/s with progressively improved persistent/direct-cache variants, with the aged-frequency experiment rejected after a 4.27 t/s regression.

Dependency

This is the persistent expert-cache/direct-dispatch feature only and is intentionally stacked on the preceding two CUDA changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant