-
Notifications
You must be signed in to change notification settings - Fork 1.9k
Pull requests: antirez/ds4
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
cuda: accelerate long-context B1 indexer and 1M HCA
#763
opened Aug 10, 2026 by
wangshang23
Loading…
metal: accelerate M5 Max indexed prefill
#758
opened Aug 9, 2026 by
rinaldofesta
Contributor
Loading…
server: anchor dsml_attr attribute match to a name boundary
#757
opened Aug 9, 2026 by
Flor1an-B
Loading…
tests: measure batched-verify greedy divergence rate, not just worst gap
#756
opened Aug 9, 2026 by
Flor1an-B
Loading…
CUDA: add DGX Spark network expert/tensor parallelism
#754
opened Aug 9, 2026 by
shankinson
Loading…
server: resolve Responses response ids to KV prefixes
#752
opened Aug 9, 2026 by
perfloop-agent
Loading…
DSpark: bitwise first-divergence diagnostics for batched vs sequential decode
#749
opened Aug 8, 2026 by
raveonetter
Loading…
cuda: give the token-tile ldmatrix its shared state space
#748
opened Aug 8, 2026 by
pmasala
Loading…
cuda: keep the raw MMQ MoE tier off streamed routed experts
#747
opened Aug 8, 2026 by
pmasala
Loading…
metal: add optional polled release fence for the TP gate (~23% decode gain with
--tensor-parallel)
#743
opened Aug 7, 2026 by
ryan5rdx
Loading…
docs: add on-disk sizes for GLM 5.2 model variants
#728
opened Aug 6, 2026 by
vincenzopalazzo
Loading…
2 of 3 tasks
server: keep live KV reusable when clients strip transient metadata blocks
#727
opened Aug 6, 2026 by
Flor1an-B
Loading…
ssd: enforce streaming cache floor at one-prefill minimum
#725
opened Aug 6, 2026 by
bestbug456
Loading…
speed-bench: add MacBook Pro M4 Max q2 sweep
#723
opened Aug 6, 2026 by
datanerdie
Contributor
Loading…
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.