Skip to content

Add Clef as a local System One referee, and pin llama.cpp b11485 - #4

Draft
markelphoenix wants to merge 14 commits into
mainfrom
cursor/clef-system-one-cb42
Draft

markelphoenix wants to merge 14 commits into
mainfrom
cursor/clef-system-one-cb42

Conversation

@markelphoenix

@markelphoenix markelphoenix commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

This is not legal advice. The notices in this pull request are plain-language warnings. They do not create a complete EULA, and they should not be treated as sufficient legal protection. A lawyer should review the items listed under "Remains for human or legal review" before this is relied on for a store page or a public release.

Do not merge from this draft.

Screenshots below are from an earlier round. They still show a tokens/s figure. The menu now shows estimated seconds per decision. This environment has no Tk, so they are not a screenshot of the game window. The pin in these pictures is llama.cpp b11485.

System One on a 32 GB RTX 5090, previous round. The speed line now says seconds per decision.

System One on an 8 GB laptop GPU. Clef-flash Q4_K_M is recommended. Full Clef is a snug fit.

Installed engine b11100 is too old for Clef. The menu offers to download llama.cpp b11485.

Steam Hardware Survey

The fit is checked on simulated machines from the common Steam range, not only the RTX 5090 test PC. Each row locks the story model and the System One referee. Download size plus 1 GB of spare disk stays inside the free disk in the fixture (200 GB). Estimated memory stays inside the planner's budget. That budget already keeps 0.8 GB of video memory spare, subtracts memory other programs are using, and keeps an extra 3 GB spare for Clef on Windows only. Story models do not take that extra cut.

A recommended story is at least 3 tokens/s for a whole turn and is not labelled "very slow". The speed word matches the raw tokens/s estimate. A local referee's estimated seconds stay under the timeout for where it runs: 30 s on the GPU or in one shared-memory pool, 120 s for a GPU+RAM split, 180 s on the CPU.

"Memory" is need / budget. On a discrete GPU the budget is usable video memory. On a unified-memory machine it is a share of system RAM (65% at 8 GB, 70% through 48 GB, 75% at 64 GB, 80% at 128 GB), and that share is not added to RAM again. On the CPU, and for the RAM side of a split, it is spare system RAM. A split also says what share of the layers stays on the card.

Windows 10 and Windows 11 use the same fit rules. Both are in the matrix so a version string cannot silently change the pick. AMD and Intel Arc rows are the Vulkan path: no NVML and no nvidia-smi. Apple Silicon is included because the game ships a Metal build. The later rows add more Mac sizes, RTX Spark and DGX Spark, Ryzen AI Max, and a Lunar Lake laptop. Those budgets are a share of system RAM.

The rows that were already in this table are unchanged by the unknown-memory margin. Each survey card states its in-use figure, including zero, so the 0.8 GB reserve still applies. A probe that never read in-use memory is a separate case and keeps about 2 GB spare.

Picks that were wrong, and what they are now

  1. Optimus. A laptop whose built-in chip reported 16 GB of shared memory was planned on that chip, and the story became a 20B model on the fake video memory. The discrete card wins even when the built-in chip's number is larger. The GTX 1650 row and the same laptop now both get Qwen3 4B Q4_K_M, mostly on the 1650.
  2. Steam Deck. "AMD Custom GPU 0405" (LCD Aerith, OLED Sephiroth) was treated as an 8 GB graphics card, so the story was Qwen3 8B on the GPU. That chip shares the 16 GB with the processor. The plan is now Qwen3 4B Q4_K_M on the CPU, same shape as a machine whose built-in graphics report no private video memory.
  3. GTX 1650 with 1.5 GB already in use. The story was Mistral 7B on the CPU at about 3.8 tokens/s. A 4B model that keeps about 45% of its layers on the card is about 10.7 tokens/s, so that is the pick. A fast split still loses to a settled model that is already at 8 tokens/s or more, so a thin split does not jump the queue on a card that can hold a comfortable model.
  4. Windows RTX 5090 with Qwen3 32B loaded. Full Clef only fits on the CPU (about 28 s). Clef-flash keeps about 59% of its layers on the card and returns in about 2.2 s, so flash is the referee. The menu says that share, and that the rest sits in system RAM. A 24 GB card cannot put either Clef file on the GPU beside that story. Both are on the CPU, and Clef Q4 is more than twice as slow as Clef-flash, so the pick is Clef-flash (about 9 s on 16 cores, about 12 s on 12 cores). Both stay under the 180 s CPU timeout.

Expected picks

Machine Story Where Download Memory Speed Referee Wait
GTX 1650 4 GB, Windows 10, 16 GB RAM Qwen3 4B Q4_K_M ok, partial 2.50 GB 3.8 / 15.7 GB fast, 17.8 tok/s Clef-flash Q4_K_M, ok, CPU 24.7 s / 180 s
RTX 2060 6 GB, Windows 10, 16 GB RAM Qwen3 8B Q4_K_M ok, partial 5.03 GB 6.1 / 17.7 GB fast, 14.9 tok/s story model decides —
RTX 3050 6 GB, Windows 11, 16 GB RAM Qwen3 8B Q4_K_M ok, partial 5.03 GB 6.1 / 17.7 GB usable, 12.6 tok/s story model decides —
RTX 3060 Ti 8 GB, Windows 11, 16 GB RAM Qwen3 8B IQ4_XS ok, GPU 4.59 GB 5.7 / 7.2 GB fast, 58.6 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RTX 4060 8 GB, Windows 11, 16 GB RAM Qwen3 8B IQ4_XS ok, GPU 4.59 GB 5.7 / 7.2 GB fast, 36.8 tok/s Clef-flash Q4_K_M, great, CPU 24.7 s / 180 s
RTX 4060 Laptop 8 GB, Windows 11, 16 GB RAM Qwen3 8B IQ4_XS ok, GPU 4.59 GB 5.7 / 7.2 GB fast, 34.7 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RTX 3060 12 GB, Windows 11, 32 GB RAM Qwen3 14B IQ4_XS ok, GPU 8.18 GB 9.1 / 11.2 GB fast, 28.0 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RTX 3060 12 GB, Linux, 32 GB RAM Qwen3 14B IQ4_XS ok, GPU 8.18 GB 9.1 / 11.2 GB fast, 28.0 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RTX 4070 12 GB, Windows 11, 32 GB RAM Qwen3 14B IQ4_XS ok, GPU 8.18 GB 9.1 / 11.2 GB fast, 38.6 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RTX 4060 Ti 16 GB, Windows 11, 32 GB RAM Qwen3 14B Q4_K_M ok, GPU 9.00 GB 9.9 / 15.2 GB fast, 20.6 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RTX 4080 16 GB, Windows 11, 32 GB RAM Qwen3 14B Q6_K ok, GPU 12.10 GB 12.8 / 15.2 GB fast, 37.4 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RTX 3090 24 GB, Windows 10, 64 GB RAM Qwen3 32B IQ4_XS ok, GPU 17.90 GB 18.6 / 23.2 GB fast, 33.2 tok/s Clef-flash Q4_K_M, great, CPU 12.4 s / 180 s
RTX 4090 24 GB, Windows 11, 64 GB RAM Qwen3 32B IQ4_XS ok, GPU 17.90 GB 18.6 / 23.2 GB fast, 35.6 tok/s Clef-flash Q4_K_M, great, CPU 9.3 s / 180 s
RTX 5090 32 GB, Windows 11, 64 GB RAM Qwen3 32B Q5_K_M ok, GPU 23.20 GB 23.5 / 31.2 GB fast, 48.1 tok/s Clef-flash Q4_K_M, ok, partial (~59% on card) 2.2 s / 120 s
Dual RTX 3060 12 GB, Windows 11, 32 GB RAM Qwen3 32B IQ4_XS ok, GPU 17.90 GB 18.6 / 22.4 GB usable, 13.1 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RX 6600 8 GB, Windows 11, Vulkan, 16 GB RAM Qwen3 8B IQ4_XS ok, GPU 4.59 GB 5.7 / 7.2 GB fast, 30.6 tok/s Clef-flash Q4_K_M, great, CPU 24.7 s / 180 s
RX 7800 XT 16 GB, Linux, Vulkan, 32 GB RAM Qwen3 14B Q6_K ok, GPU 12.10 GB 12.8 / 15.2 GB fast, 32.8 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
Intel Arc A770 16 GB, Windows 11, Vulkan, 32 GB RAM Qwen3 14B Q6_K ok, GPU 12.10 GB 12.8 / 15.2 GB fast, 29.5 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
Intel Iris Xe, Windows 11, 16 GB RAM Qwen3 1.7B Q4_K_M great, CPU 1.11 GB 2.1 / 12.5 GB fast, 13.6 tok/s Clef-flash Q4_K_M, ok, CPU 37.1 s / 180 s
AMD Radeon 780M, Windows 11, 16 GB RAM Qwen3 4B Q4_K_M great, CPU 2.50 GB 3.5 / 12.5 GB usable, 11.1 tok/s Clef-flash Q4_K_M, ok, CPU 18.6 s / 180 s
CPU only, Windows 10, 8 GB RAM Qwen3 1.7B Q4_K_M great, CPU 1.11 GB 2.1 / 4.5 GB usable, 8.7 tok/s story model decides —
CPU only, Windows 11, 16 GB RAM Qwen3 1.7B Q4_K_M great, CPU 1.11 GB 2.1 / 12.5 GB fast, 13.6 tok/s Clef-flash Q4_K_M, ok, CPU 24.7 s / 180 s
CPU only, Linux, 32 GB RAM Qwen3 1.7B Q5_K_M great, CPU 1.26 GB 2.2 / 29.5 GB fast, 13.7 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
Optimus: Iris Xe + RTX 4060 Laptop 8 GB Qwen3 8B IQ4_XS ok, GPU 4.59 GB 5.7 / 7.2 GB fast, 34.7 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
Optimus: Iris Xe reported as 16 GB + GTX 1650 4 GB Qwen3 4B Q4_K_M ok, partial 2.50 GB 3.8 / 15.7 GB fast, 17.8 tok/s Clef-flash Q4_K_M, ok, CPU 24.7 s / 180 s
Steam Deck, Linux, 16 GB shared, Custom GPU 0405 Qwen3 4B Q4_K_M great, CPU 2.50 GB 3.5 / 13.5 GB fast, 13.9 tok/s Clef-flash Q4_K_M, ok, CPU 37.1 s / 180 s
Apple M2, 16 GB unified (about 11.2 GB GPU share) Mistral 7B Q4_K_M great, unified 4.37 GB 5.5 / 11.2 GB usable, 11.6 tok/s Clef-flash Q4_K_M, tight, partial (~72% on the GPU share) 1.1 s / 120 s
Apple M2 Pro, 16 GB unified (about 11.2 GB GPU share) Qwen3 14B IQ4_XS ok, unified 8.18 GB 9.1 / 11.2 GB usable, 12.5 tok/s story model decides —
RTX 3060 12 GB with 2.5 GB already in use, 32 GB RAM Qwen3 8B Q5_K_M ok, GPU 5.85 GB 6.9 / 8.7 GB fast, 38.3 tok/s Clef-flash Q4_K_M, great, CPU 18.6 s / 180 s
RTX 4060 8 GB with 2 GB already in use, 16 GB RAM Qwen3 8B Q4_K_M ok, partial 5.03 GB 6.1 / 17.7 GB fast, 14.6 tok/s story model decides —
GTX 1650 4 GB with 1.5 GB already in use, 16 GB RAM Qwen3 4B Q4_K_M ok, partial (~45% on card) 2.50 GB 3.8 / 14.2 GB usable, 10.7 tok/s Clef-flash Q4_K_M, ok, CPU 24.7 s / 180 s
Apple M1, 8 GB unified (about 5.2 GB GPU share) Qwen3 4B Q4_K_M ok, unified 2.50 GB 3.8 / 5.2 GB usable, 13.3 tok/s story model decides —
Apple M4, 24 GB unified (about 16.8 GB GPU share) gpt-oss 20B MXFP4 ok, unified 12.10 GB 12.4 / 16.8 GB fast, 23.3 tok/s Clef-flash Q4_K_M, tight, partial (~55% on the GPU share) 4.1 s / 120 s
Apple M4 Pro, 36 GB unified (about 25.2 GB GPU share) Qwen3 14B Q4_K_M great, unified 9.00 GB 9.9 / 25.2 GB fast, 15.2 tok/s Clef-flash Q8_0, ok, unified 0.3 s / 30 s
Apple M4 Max, 64 GB unified (about 48 GB GPU share) Qwen3 32B Q4_K_M great, unified 19.80 GB 20.3 / 48.0 GB fast, 14.0 tok/s Clef Q4_K_M, ok, unified 0.6 s / 30 s
Apple M3 Ultra, 128 GB unified (about 102.4 GB GPU share) Qwen3 32B Q6_K great, unified 26.90 GB 27.0 / 102.4 GB fast, 15.3 tok/s Clef Q8_0, great, unified 0.9 s / 30 s
RTX Spark, Windows 11 on Arm, 32 GB (about 22.4 GB share) Qwen3 14B Q5_K_M great, unified 10.50 GB 11.3 / 22.4 GB fast, 14.5 tok/s Clef-flash Q4_K_M, tight, unified 0.2 s / 30 s
RTX Spark, Windows 11 on Arm, 64 GB (about 48 GB share) Qwen3 14B Q5_K_M great, unified 10.50 GB 11.3 / 48.0 GB fast, 14.5 tok/s Clef Q4_K_M, great, unified 0.6 s / 30 s
RTX Spark, Windows 11 on Arm, 128 GB (about 102.4 GB share) Qwen3.8 27B Q4_K_M great, unified 16.46 GB 17.2 / 102.4 GB usable, 9.6 tok/s Clef Q4_K_M, great, unified 0.6 s / 30 s
DGX Spark (GB10), Linux arm64, 128 GB (about 102.4 GB share) Qwen3.8 27B Q4_K_M great, unified 16.46 GB 17.2 / 102.4 GB usable, 9.6 tok/s Clef Q4_K_M, great, unified 0.6 s / 30 s
Ryzen AI Max 8060S, Windows 11, 64 GB (about 48 GB share) Qwen3 14B Q4_K_M great, unified 9.00 GB 9.9 / 48.0 GB fast, 14.3 tok/s Clef Q4_K_M, great, unified 0.6 s / 30 s
Ryzen AI Max 8060S, Linux, 128 GB (about 102.4 GB share) Qwen3 14B Q4_K_M great, unified 9.00 GB 9.9 / 102.4 GB fast, 14.3 tok/s Clef Q4_K_M, great, unified 0.6 s / 30 s
Lunar Lake Arc 140V, Windows 11, 32 GB (about 22.4 GB share) Qwen3 30B-A3B (MoE) Q4_K_M ok, unified 18.60 GB 18.6 / 22.4 GB fast, 26.1 tok/s story model decides —
Windows, 2 GB RAM, no GPU no local model — — — — no local referee —

When local AI does not fit

  • 2 GB RAM. The story menu says the models are all too big or too slow, and offers the pretend model (mock).
  • Clef does not fit beside the story (RTX 2060 and RTX 3050 6 GB with 16 GB RAM, an 8 GB CPU, an M2 Pro whose 14B story fills the GPU share, a 4060 with 2 GB already in use, an 8 GB M1, and a Lunar Lake laptop whose 30B-A3B story fills the share). The referee menu says "Neither Clef model looks like it fits in the memory you have left" and offers "No extra referee — the story model decides". That is the story model, not the pretend model. The story is not shrunk just to make room for Clef.
  • M2 16 GB referee is tight (7.9 GB need, 8.0 GB left). It still fits, at about 1.1 s, under the 120 s split timeout. The label says tight on purpose.
  • CPU Clef at about 55 s on 8 cores is still the measured anchor for the 19.23 GB file. When Clef-flash is also on the CPU, that file is about 19 s, more than twice as quick, so flash is the referee. The 55 s figure stays under the 180 s CPU timeout.

The two 12 GB NVIDIA rows (Windows 11 and Linux) pick the same story and the same referee. Linux keeps 2.5 GB of RAM headroom and Windows keeps 3.5 GB, so the RAM budget differs (29.5 GB vs 28.5 GB) and the pick does not. The dual 3060 pools both cards: usable video memory is 2 × (12 − 0.8) GB.

Unified memory

Apple Silicon already planned inside one pool. This round uses that for the other machines whose CPU and GPU share RAM, and it tightens the Apple share on small and large Macs.

The GPU budget is a share of system RAM: 65% at 8 GB, 70% through 48 GB, 75% at 64 GB, and 80% at 128 GB. That percentage is the headroom. The fit does not also subtract the OS reserve from the share, and it does not add the share to RAM a second time. A BIOS "dedicated VRAM" number is not the budget by itself. When DXGI reports dedicated plus shared, the budget is the sum, capped at the share. A carve-out with no shared figure uses the share and ignores the carve-out. nvidia-smi reporting N/A, or reporting the whole RAM stick, does not become a second pool. Used memory against a whole-pool reading is scaled down to the share.

Steam Deck's Custom GPU 0405 stays shared RAM with a CPU plan. A discrete card still wins over a built-in chip, including a named NVIDIA card whose size was never read. A lone Iris Xe, UHD, Phoenix, or 780M detected live is now that RAM share (Vulkan when the engine can use it). The Iris Xe and 780M rows already in the table still pass zero video memory, so those CPU picks stay locked.

Speed for RTX Spark, DGX Spark (GB10), and GB300 uses an LPDDR5X-9400 class guess, 301 GB/s, not a discrete-GPU table. Apple chips keep their M1–M4 Pro/Max/Ultra bandwidth tiers, and the engine is Metal. Strix Halo (Radeon 8060S) stays at 256 GB/s. Lunar Lake Arc 140V stays at 136 GB/s. Strix Halo and Lunar Lake prefer Vulkan.

Windows on Arm with an NVIDIA GPU plans CUDA 13 arm64 when the driver is 580 or newer (or unknown), then Vulkan arm64, then CPU arm64. An older driver starts at Vulkan arm64. Linux arm64 with NVIDIA is CUDA 13 arm64, then Vulkan arm64 when the loader is present, then CPU arm64. The setup menu names that build and says an x64 engine is not run under emulation. select_assets matches the host architecture, so an arm64 plan cannot return an x64 archive name.

b11485 publishes those arm64 archives. The pin already had their SHA-256 lines, and this round checked them against the release API. They match. There is no CUDA 12 arm64 build.

  • Windows: llama-b11485-bin-win-cuda-13.4-arm64.zip and cudart-llama-bin-win-cuda-13.4-arm64.zip
  • Ubuntu: llama-b11485-bin-ubuntu-cuda-13.4-arm64.tar.gz and cudart-llama-b11485-bin-ubuntu-cuda-13.4-arm64.tar.gz
  • Vulkan: llama-b11485-bin-win-vulkan-arm64.zip and llama-b11485-bin-ubuntu-vulkan-arm64.tar.gz
  • CPU: llama-b11485-bin-win-cpu-arm64.zip and llama-b11485-bin-ubuntu-arm64.tar.gz
  • macOS arm64 was already pinned (the Metal build)

The survey rows use the share (22.4 GB on 32 GB, 48 GB on 64 GB, 102.4 GB on 128 GB). Two detection cases are tighter than that share. A 128 GB Strix Halo whose DXGI numbers are 64 GB dedicated and 32 GB shared is budgeted at 96 GB. A 32 GB Lunar Lake whose DXGI numbers are 0.125 GB dedicated and 15.5 GB shared is budgeted at 15.6 GB. The matrix row for Lunar Lake is the 22.4 GB share, which is what the fit uses when Windows does not report a shared figure beside a tiny carve-out.

GB300 uses the same budget and the same 301 GB/s guess as DGX Spark. The survey table does not add a separate GB300 row.

Model discovery (already on this branch)

Three model-discovery bugs, read against commit d62ef5f. The six hardware-test fixes from the previous round stay as they are. Abliterated and uncensored repos are still rejected. The family-friendly filter and the Steam disclosure are unchanged. Underdog Saluki is not a catalog entry; its byte and parameter counts are only a test of the bits-per-weight math. The GSQ-RCO repo is not a catalog entry either.

  1. Qwen3.8 27B GGUF repos can be found. Hub search used filter="gguf" with pipeline_tag="text-generation". The major Qwen3.8-27B GGUF repos are tagged image-text-to-text or have no pipeline tag (unsloth, ggml-org, bartowski, lmstudio-community). Each publisher search now also asks for image-text-to-text and for repos with no pipeline tag, then keeps text-generation, image-text-to-text, and untagged results. They can still be used for text only. mmproj files are still skipped. License, publisher, and specialist checks are the same, so a vision-named, speech, coder, or abliterated repo is still dropped.

    unsloth/Qwen3.8-27B-GGUF is a seed: Apache-2.0, base Qwen/Qwen3.8-27B, architecture qwen35 (llama.cpp b11485 loads it), 27.32B parameters, KV shape 64 layers × 4 KV heads × 256. The published files are Unsloth's UD variants; the seed stores those sizes under the plain quant names (Q8_0 29.05 GB, Q6_K 21.98, Q5_K_M 19.77, Q4_K_M 16.46, IQ4_XS 14.25). There is no Q3_K_M file in that repo, so the seed does not invent one. A built game that can load the curated list treats qwen35 as known, so the seed is offered.

    On a 32 GB card with 64 GB RAM the recommendation stays Qwen3 32B (on the card). Qwen3.8 27B at Q5_K_M is close behind, also on the card, with a lower score. On a 12 GB card both spill, the 27B spill scores higher than the 32B spill, and the pick is Qwen3 14B, which still fits on the card.

  2. A LoRA adapter is not a full model. parse_quant and pick_gguf_file skip any path containing lora (case-insensitive), next to the existing mmproj skip. A path like lora/NAME-LoRA-r320-F16.gguf is no longer returned as an F16 weight.

  3. Quality follows the file, not a misread label. parse_quant accepts IQn-mix and IQn_mix. entry_from_hub keeps that file. When the file size and the parameter count are known, bits per weight are file bytes × 8 ÷ parameter count. The label is used only when the size is unknown. An unknown label scores 0.55, not the old 0.9. Saluki's published figures (7,898,369,152 bytes, 26,895,998,464 parameters) are about 2.35 bits per weight and stay a last resort. ISTA-DASLab's GSQ-RCO file named IQ3_S (11,771,546,784 bytes, same parameter count, about 3.50 bits) is the low rung, so a comfortable fit is not forced into "heavily compressed."

    A known quant is not scored above its name when the file is only a little larger than the label (tokenizer and metadata). Clef's Q4_K_M stays on the 4-bit rung, so an 8 GB laptop still gets Clef-flash, and a Mac still prefers the quant that fits on the GPU. A last-resort name whose file is really at least 3.5 bits is promoted. A file that measures worse than its name uses the measured score, so a 2.35-bit file cannot hide behind a Q4 label.

Hardware round (already on this branch)

Lux's Windows 11 machine (Ryzen 7 9850X3D, 64 GB RAM, RTX 5090 32 GB, driver 617.14, Clef Q4_K_M, llama.cpp b11485 CUDA 13) scored 5/5 prompts at context 4096 in 0.35–0.88 s and used 19.1 GiB. Context 16384 also worked. Six bugs from that run are fixed here.

  1. Tests no longer touch a real GPU. The specs fixture and an autouse guard stub NVML, the Windows registry, and DXGI. If a test still reaches the real probe, it raises before any DLL loads. On that PC, four specs tests failed because the NVML fallback loaded nvml.dll. Those tests pass with the stubs. On that round, tests/test_specs.py was 60/60 on this Linux VM. The survey adds the Steam Deck case on top of that.

  2. The physical batch is capped at 4,096. -c can stay large. -b and -ub are min(context, 4096). Matching them to a 65,536 context made the compute buffer 26,519 MiB and CUDA failed. The same 65,536 context with a 4,096 batch stayed on the GPU at about 18.7 GB. The request builder trims a prompt that would exceed the batch, so the server is not asked to score something the batch cannot hold. Clef's context cache measured 0 MiB, so the planner no longer grows a KV estimate with -c. The compute buffer in the estimate stays about 1.3 GB, the size measured at batch 4,096.

  3. The referee waits longer off the GPU. A full-GPU call still times out at 30 s. A GPU+RAM split waits 120 s. A CPU call waits 180 s. The measured CPU rounds were 47–65 s, so every one of them used to fall back to the local judge. A local timeout says the hardware is busy. The hosted Jev message is unchanged.

  4. Free memory, not the card's total. nvidia-smi / NVML in-use memory is subtracted from every menu fit budget. Windows adds that same in-use figure to --fit-target, because CUDA there had reported about 30,991 MiB free while other programs were using 1–6 GB. Linux --fit-target stays the 0.8 GiB reserve: CUDA's free figure on Linux already excludes other programs, so adding them again would keep that memory spare twice. "Needs ~X of Y GB" uses the free figure. An extra 3 GB on Windows applies to Clef's budget and Clef's --fit-target only, because that run used about 2.9 GB more than llama.cpp projected. Story models do not take that cut: a 6 GB Windows card would otherwise miss the 4B story model the store page promises. A Windows story launch still adds in-use memory to --fit-target. With nothing in use and no specs, the documented target stays 819 MiB.

  5. More memory pressure does not pick a bigger referee. A GPU+RAM split is a worse home than a model that still fits wholly on the card. With Qwen3 8B loaded, full Clef was a snug full-GPU fit and Clef-flash won. With Qwen3 14B loaded, full Clef became a 75% split and used to win instead. The larger model no longer wins that comparison. The menu shows estimated seconds per decision (about 0.6 s fully on the card, about 1 s at a 75% split, about a minute on 8 CPU cores for Clef Q4), not tokens/s. A real 75% split (-ngl 48) took about 1.0 s per round. The story menu's tokens/s column is unchanged.

  6. The engine pin covers every archive the Clef installer can download. SHA-256 lines were added for the CUDA 13 and CUDA 12 builds, the CUDA runtime zips, and the Vulkan and CPU archives for each OS that path can select. Windows CUDA runtime zips omit the build tag in the filename; the pin parser accepts that. install_tagged_release checks the pin before download and fails closed when the pin has no digest or GitHub's digest differs. The five game-build digests are unchanged. The new lines are the sha256 values GitHub publishes on the b11485 release. Those zip files were not downloaded again on this machine.

(a) What changed, including earlier rounds

Clef is Cloudflare's open-weight decision model. It scores labelled answers in one forward pass and does not write free-form story text, so it is a System One referee, not a storyteller. It is not in MODEL_CATALOG.

  • SYSTEM_ONE_CATALOG: Clef (27.02B, ggml-org/Clef-GGUF) and Clef-flash (9.08B, ggml-org/Clef-Flash-GGUF). Quant sizes are the published GGUF byte lengths. License is Apache-2.0. Requested context is 4,096 tokens. Published native context is 65,536. Architecture is clef.
  • recommend_system_one prefers the larger comfortable model on the best home. A split does not count as "on the card." A card-resident fit at least four times quicker, in seconds per decision, replaces a larger CPU fit. A partial split counts as card-resident for that comparison.
  • The story recommender can pick a fast GPU+RAM split when every settled model has fallen below 8 tokens/s. A settled model that is already that quick still wins.
  • The primary card is the discrete GPU with the most video memory when one is present. A built-in chip is skipped even when its shared memory is reported as video memory, as long as that discrete card is there. A lone built-in chip, or RTX Spark / GB10 / GB300, is one pool: a share of system RAM. Steam Deck's "Custom GPU 0405" stays a CPU plan.
  • Local Clef passes --no-repack, --fit on with the margin above, and -b/-ub capped at 4,096. CPU launches pass --fit off, --device none, and -ngl 0. A GPU or split launch also passes --device for the cards left after the story model is reserved.
  • The engine pin is b11485. An older installed engine is offered that download when downloads are allowed. The weight file is still refused until the engine can load it.
  • Confirm screen: hardware strain, "AS IS" / no warranty, and "AI output can be wrong. You are responsible for how you use it; it is not advice."
  • Story catalog now includes Qwen3.8 27B (unsloth/Qwen3.8-27B-GGUF, Apache-2.0, architecture qwen35). Discovery can also see image-text-to-text and untagged GGUF repos.

(b) Evaluation notes that this round changes

Finding What we did
A laptop's built-in chip with a large shared-memory number beat the discrete card. Skip integrated GPUs when choosing the primary card.
Steam Deck's Custom GPU 0405 looked like an 8 GB discrete GPU. Treat Aerith / Sephiroth / Custom GPU 0405 as shared memory.
A 4 GB card with memory in use recommended a slow CPU 7B over a faster 4B split. A great/ok fit at 8 tokens/s or more can win even when the card holds less than 75% of the layers.
A Windows 5090 with a 32B story recommended full Clef on the CPU (~28 s) over Clef-flash on the card (~2.2 s). A partial referee at least four times quicker replaces the CPU pool. The reason names the share that stays on the card.
Qwen3.8-27B GGUF repos were invisible because they are not tagged text-generation. Search also accepts image-text-to-text and untagged GGUF repos. Seed the Unsloth Apache-2.0 repo. A comfortable 32B stays the 5090 pick.
lora/...-F16.gguf was chosen as a full F16 weight. Skip any path containing lora.
IQ2-mix was dropped, and a label could overrate a 2-bit file or underrate a 3.5-bit file named IQ3_S. Parse IQn-mix. Use measured bits per weight when size and parameter count are known. Unknown labels score 0.55. A known name is not promoted by tokenizer overhead.
Tests on an NVIDIA PC loaded real nvml.dll. Stubs plus a pytest guard.
-b/-ub equal to a large context filled the card. Cap at 4,096. Trim the request to the batch. Planner keeps Clef's KV at 0 and the compute buffer at the measured 1.3 GB.
30 s timeout fell through on CPU. 30 / 120 / 180 s by placement. Local wording names hardware speed.
--fit-target 819 ignored memory already in use, and the projection was short on Windows. The menu budget subtracts in-use memory on every OS. Windows --fit-target adds it; Linux stays at 819 MiB. Add 3 GB for Clef on Windows only.
A tighter machine recommended the larger referee, and the speed was tokens/s. A split is never the same home as a full-GPU fit. Show seconds per decision.
The Clef installer trusted only GitHub's checksum. Pin every archive that path can download, and fail closed.

Story-catalog spot checks that already looked right were left alone. The story recommender's 8 tokens/s floor for the first tier is still intentional. A remembered Clef choice still asks again instead of auto-starting. LEARN's hosted-Jev pages still say Jev.

(c) Liability checklist

This is not legal advice.

Found, then kept

  • MIT license with an "AS IS", no-warranty disclaimer. That is a software license disclaimer, not a EULA covering hardware damage, AI output, or third-party terms.
  • README already said estimates can be wrong, not affiliated, trademarks, the player is responsible for model licenses, and Jev may cost money.
  • Privacy section lists huggingface.co, api.github.com, api.typesafe.ai, and says there is no telemetry.
  • Store disclosure still says the game creates no AI images. The referee does not download the vision file. The family-friendly filter is unchanged.

Added or reworded

  • Confirm screen and Clef download path: hardware strain, "AS IS" without warranty, and "AI output can be wrong. You are responsible for how you use it; it is not advice."
  • README Clef section now describes seconds per decision, the 4,096 batch cap, the longer CPU wait, and free memory rather than the card's total. The hardware-check step says a laptop is planned on the discrete card, and that Steam Deck's 16 GB is shared.
  • README hardware-check step also says the engine is started on the cards that plan counted, and that the referee is started on the card its own plan counted. The Clef memory paragraph says an unread in-use figure keeps about 2 GB spare instead of 0.8 GB.
  • README hardware-check step now also says RTX Spark, DGX Spark, GB300, Ryzen AI Max, Lunar Lake, and a generic iGPU are one share of RAM. The memory-available bullet lists the 65/70/75/80% shares and says that share is not added to RAM again. LEARN and ARCHITECTURE describe the same pool, and Windows on ARM is CUDA 13 arm64, then Vulkan arm64, then CPU arm64.
  • NOTICE: Apache-2.0 attribution for Clef and Clef-flash. The llama.cpp row names pin b11485.
  • Steam disclosure names local Clef, b11371 text support, the b11485 pin, text-only referee use, and the upgrade offer.
  • LEARN and ARCHITECTURE describe the wider GGUF search (text-generation, image-text-to-text, and untagged) and that mmproj files are still skipped. ARCHITECTURE also says the primary card is the discrete GPU, and that Steam Deck's Custom GPU 0405 shares its memory with the CPU.

Remains for human or legal review

  • Whether the MIT disclaimer plus these plain-language notices are enough for a commercial Steam release, including hardware damage, heat, and power draw.
  • Third-party acceptable-use and trademark terms: Cloudflare, Qwen, Hugging Face, TypeSafe, ggml / llama.cpp, Ollama, and the Steam Subscriber Agreement.
  • A privacy policy if the game is ever hosted, or if telemetry is added.
  • Do not describe these notices as legally sufficient.

(d) Sources

Workers AI was not called. No third-party VRAM calculator was used. The new pin lines were copied from GitHub's published digests, not hashed again from the zip bytes on this machine. The arm64 CUDA 13.4, Vulkan, and CPU archives for Windows and Ubuntu were already in that pin. This round checked those digests against the b11485 release API again. They match. There is no CUDA 12 arm64 archive. Those zip bytes were not downloaded or run. The original five game-build archives were hashed from the downloaded bytes in an earlier round and already matched those digests. Qwen3.8 file sizes and the GGUF parameter count were read from the Hub listing. Saluki and the GSQ-RCO IQ3_S file were used only as test figures; neither repo was added. The survey machines are simulated specs. They were not measured on those cards.

Clef on the earlier CPU smoke

Clef-flash Q4_K_M did run here, on the Ubuntu CPU build, before the hardware round. Three geese prompts scored in 54–60 s with output_tokens: 0. That smoke used -b/-ub 2048 at context 2048. The game now caps the batch at 4,096 instead of matching whatever context was requested.

Planned card, unknown VRAM, and a bad file

These checks were reimplemented in this repository. No code, patterns, or tests were copied from the AGPL-3.0 project that suggested them.

  1. The engine stays on the planned cards. --list-devices is matched to the cards the fit already pooled (the primary card, plus other cards from that vendor with at least 4 GB). The launch passes --device for those rows. Two NVIDIA cards stay CUDA0,CUDA1, with CUDA_DEVICE_ORDER=PCI_BUS_ID and CUDA_VISIBLE_DEVICES=0,1. A Vulkan listing that also shows an Intel or AMD built-in chip passes only the discrete card. A lone CUDA1 is renamed to CUDA0 once it is the only visible device. The list-devices probe sets the PCI order for NVIDIA and does not set CUDA_VISIBLE_DEVICES. The Clef process uses that listing when the engine file is the same, or reads --list-devices once for its own build. It is pinned to the cards left after the story model's video memory is reserved. That can be a different card: a Vulkan build that sees an RTX 4090 and an RX 7800 XT starts the story on the 4090 and, once that card is reserved, starts Clef on the 7800 XT. Two identical cards keep their list position, so a referee planned on the second card is CUDA_VISIBLE_DEVICES=1 (the engine sees it as CUDA0). A built-in chip in the same listing is still left out.
  2. Unknown in-use memory keeps about 2 GiB. vram_used_known is false when nvidia-smi had no used column, NVML returned no used figure, Linux AMD had no mem_info_vram_used, or the size came from the Windows registry or DXGI. The menu fit then keeps 2 GiB instead of 0.8 GiB plus a measured in-use value. On Windows, --fit-target does the same, and the Clef extra 3 GB is still added on top. On Linux, --fit-target stays 819 MiB, because there is no separate CUDA-free reading and the free figure already excludes other programs. Linux AMD reads mem_info_vram_used and clamps it to the total.
  3. A bad GGUF is refused before the engine is stopped or started. The file must start with GGUF and a version from 1 to 3. The message tells the player to delete it and download it again, and says the engine was not started. The same check runs for an mmproj path when one is set. This game does not download a projector. Clef's launch runs that check after the download returns and before the process starts, and raises LocalClefUnavailable with the same refusal.
  4. A corrupt file is not a GPU failure. gguf_init_from_ (any reader), wrong number of tensors, and failed to open gguf are classified as the model file before a CUDA or Vulkan error line. A CUDA out-of-memory line that also says failed to load model stays a GPU error. Those markers match the public llama.cpp phrases. This machine did not replay a corrupt file on b11485.
  5. The missing CPU instruction is named when the OS probe looks for it: AVX, AVX2, and SSE4.2 on Windows; those plus FMA and F16C on macOS; those plus BMI2 on Linux (/proc/cpuinfo now records bmi2). An empty flag list or a NEON-only chip keeps the generic sentence. Windows has no standard check for FMA, F16C, or BMI2, so those are not reported as missing there.
  6. After 60 seconds, a CPU load says the model is fully on the CPU, so loading takes longer. A split says the model is partly on the CPU, so loading takes longer. A full GPU load keeps the original spinner.
  7. Small probe fixes. Display-driver registry subkeys are every four-digit name, not only 0000–0015. nvidia-smi drops a card whose total is missing or not positive, and clamps used memory into the range from 0 to that total.

Nothing in those seven items was already implemented. The survey table did not move for that round.

ALUCARD rerun

Lux reran the game at 9a092bc on the Windows RTX 5090 (ALUCARD). That run passed: Windows pytest 3779 passed and 0 failed; Clef at 4096, 16384, and 65536 context all launched with -b/-ub 4096, peaked at 21,746 MiB, and scored in 0.3–0.8 s; the device pin gave CUDA0 and left out the AMD iGPU; the GGUF check works; checksums fail closed; and with VRAM in use the planner picked Clef-flash at 86% on the GPU, with both models peaking at 27.7 of 32.6 GB. Three gaps from it are fixed here. The unified-memory commit did not move the 5090 survey row: with nothing in use, the story is still Qwen3 32B Q5 on the GPU and the referee is still Clef-flash Q4 with about 59% of its layers on the card (about 2.2 s / 120 s).

  1. The Clef launch checks the GGUF header. download_gguf returns, then gguf_file_problem runs, and a bad file raises LocalClefUnavailable with the same refusal the story engine uses. The process is not started.
  2. Memory already in use pushes that partial off the card. About 2150 MiB in use rounds to 2.1 GiB. With the 0.8 GiB reserve and the 3 GB Windows Clef margin, --fit-target is 6042 MiB. Qwen3 32B Q5 still sits fully on the GPU at about 23.5 GB, and both Clef files then only fit on the CPU (about 55 s and about 19 s on 8 cores). A partial-GPU Clef-flash still beats a CPU referee when the card can hold one. When both are on the CPU, the faster file wins if the larger is more than twice as slow or over about 30 s. The same rule moves the other CPU-only Clef rows in the table to Clef-flash.
  3. Linux --fit-target does not add in-use memory. Windows CUDA's free figure ignores other programs, so that launch still adds them. Linux CUDA's free figure already excludes them. The menu budget still subtracts in-use memory on both. There is no stored CUDA-free field, so the split is the OS.

Tests

TERM=dumb PYTHONPATH=src python3 -m pytest: 3768 passed, 60 skipped, 0 failed (33.11s). The previous full suite was 3766 passed, 60 skipped. The 5090 survey row is unchanged: Clef-flash Q4 with about 59% of its layers on the card, about 2.2 s / 120 s. The unified-memory commit did not move that row. Twelve rows that had full Clef on the CPU now pick Clef-flash on the CPU, because that file is more than twice as fast. One earlier run of the RAM-bandwidth buffer test hit the known noisy-neighbour ratio (about 2.05 against a limit of 2.0) and was not loosened.

Not verified on this machine

  • Real NVML, DXGI, or nvidia-smi used-memory on Windows. The probes are stubbed under pytest. The in-use rows set vram_used_gb on the simulated card.
  • CUDA honoring the new --fit-target, including the extra 3 GB on a Clef launch.
  • A 65,536 context with the new 4,096 batch cap. Lux already measured that combination at about 18.7 GB, all on the GPU, before this code landed.
  • Whether 180 s is enough on that 8-core CPU path beyond the 47.2, 51.4, and 64.7 s rounds already measured. The 55 s CPU figure for full Clef is the same estimate, not a new measurement. The table now picks Clef-flash on that path (about 19 s) when both files are on the CPU.
  • The ALUCARD figures above (0.3–0.8 s scores, 21,746 MiB peak, CUDA0 pin, Clef-flash at 86% with both models at 27.7 of 32.6 GB, and Windows pytest 3779 passed) are from the rerun at 9a092bc. This commit was not run on that PC.
  • The new CUDA and CUDA-runtime pin lines against the downloaded zip bytes. They are GitHub's published digests.
  • Ollama, vLLM, and LM Studio with architecture clef or qwen35.
  • A live Hub search for unsloth/Qwen3.8-27B-GGUF. The three-query search is covered by a fake Hub client.
  • A real Steam Deck, a real AMD or Arc card, or a real Optimus laptop. Those rows are simulated names, VRAM, and bandwidth guesses.
  • --device and CUDA_VISIBLE_DEVICES on a real dual-GPU or Vulkan machine, including the Clef process. The story and referee cases are simulated listings.
  • mem_info_vram_used on a real AMD GPU.
  • The exact b11485 log text for a corrupt GGUF. The classifier looks for the public phrases above.
  • RTX Spark, N1X, DGX Spark, GB300, Strix Halo, Lunar Lake, or a real Mac recommendedMaxWorkingSetSize. The shares and the 301 GB/s LPDDR5X-9400 guess are simulated.
  • nvidia-smi N/A or a whole-pool reading on Windows arm64 or Linux aarch64. The probes are stubbed.
  • DXGI dedicated-plus-shared on a Strix Halo or Lunar Lake laptop. The 96 GB and 15.6 GB cases are stubbed tuples.
  • Running the arm64 CUDA, Vulkan, or CPU binaries. The pin matches GitHub's published digests. The archives were not executed, so an x64 engine under emulation was not observed on hardware. The asset picker is what refuses an x64 name.

To show artifacts inline, enable in settings.

Open in Web Open in Cursor 

cursoragent and others added 2 commits October 8, 2026 02:14
Clef and Clef-flash are offered after the story model is ready, with published
GGUF sizes and an Apache-2.0 notice. The pinned llama.cpp build cannot load
them, so the menu explains that and does not download the weights. Hardware
matching now names multi-GPU fits honestly, and the download confirm screen
states the hardware, warranty, and output limits.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
The pin is the newest release that still publishes the five archives the
game builds use. An older installed engine is offered that download instead
of stopping. Each graphics card keeps its own memory reserve, multi-GPU
speed is a layer split, and a fast card-resident referee is not passed over
for a much slower CPU fit. The AI-output line stays objective: output can
be wrong, and the player is responsible for how they use it.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
@cursor cursor Bot changed the title Add Clef as a local System One referee, and tighten fit warnings Add Clef as a local System One referee, and pin llama.cpp b11485 Oct 8, 2026
cursoragent and others added 12 commits October 8, 2026 03:34
TERM=dumb was forcing 80 columns even when a width was set, and NO_COLOR
was stripping colour the tests asked for. The harness treats that terminal
as capable. The smoke script, which starts its own Python, gets wide
columns and no colour so its engine line stays one sentence.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
b11485 turns --fit on with a 1024 MiB device margin, which would spill
layers the menu said still fit. GPU launches now pass the menu's 819 MiB
reserve and the planned context. CPU launches pass --fit off, because the
fitter treats leftover RAM as unlimited.

Clef's head reads rows of a quantized output tensor. The default repack
buffer cannot do that and the server aborts before it listens. Embedding
mode also shrinks the batch to 512, which is smaller than a real referee
prompt. Local Clef now passes --no-repack and a physical batch equal to
the context.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
Tests stub NVML, the registry, and DXGI so they cannot load a real GPU.
Clef's physical batch stays at most 4,096 while the context stays large,
and a request that would not fit is trimmed before it is sent. The client
waits longer when the model is off the GPU, and a local timeout talks
about hardware speed. The fit budget and --fit-target subtract memory
already in use, with extra margin on Windows. A tighter machine no longer
recommends the larger referee, and the menu shows seconds per decision.
The engine pin covers every archive the Clef installer can download, and
a checksum mismatch fails closed.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
A 3 GB cut on every Windows GPU made a 6 GB card too small for the 4B
story model the store page promises. Story budgets still subtract memory
already in use. Clef's budget and --fit-target keep the extra 3 GB on
Windows, which is where the measured run overran the projection.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
Discovery now keeps image-text-to-text and untagged GGUF repos, still
skipping mmproj files and every existing license and publisher check.
unsloth/Qwen3.8-27B-GGUF is a seed so a 32 GB card has an official 27B
option; a comfortable Qwen3 32B stays the pick. LoRA paths are not
treated as full weights. When the file size and parameter count are
known, quality and fit use bytes times 8 divided by parameters, and an
unknown label scores low until the size is known.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
Measured bits still score an unknown or last-resort name, and a file
that is really more compressed than its label. A Q4_K_M that only
measures high because of the tokenizer stays on the 4-bit rung, so an
8 GB laptop still gets Clef-flash and a Mac still prefers the quant
that fits on the GPU. The built-game architecture list includes qwen35.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
A laptop's built-in chip is skipped even when it reports more memory
than the discrete card. Steam Deck's Custom GPU 0405 is shared RAM,
so the plan stays on the CPU. A fast split beats a slow CPU story, and
a partial Clef that is several times quicker replaces a CPU referee.
The survey locks a story and a referee for the common Steam tiers.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
…unch.

When in-use video memory was never read, keep about 2 GiB spare. Class a corrupt model file ahead of a GPU switch, and name the CPU instruction the probe can see is missing.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
The launch check reads the version word after the magic. A zero version is refused, so the fake model now says version 3.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
The story launch already passes --device. Clef did not, so --fit could
spread onto the story card or a built-in chip. Reuse that build's device
listing when the engine file matches, or read --list-devices once, and
pin the referee to the cards left after the story model is reserved.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
RTX Spark, DGX Spark, and GB300 use a share of system RAM, not VRAM plus
RAM, with an LPDDR5X-class speed guess. Apple Silicon keeps a
recommended-working-set share and Metal. Strix Halo and Lunar Lake use
dedicated plus shared memory, capped at that share. Windows and Linux
Arm stay on the pinned arm64 CUDA, Vulkan, or CPU build.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
The Clef launch checks the GGUF header before the process starts. On a
5090 with the 32B story loaded and about 2.1 GiB already in use, both
referees are on the CPU, and the faster file wins. Linux --fit-target
no longer adds in-use memory that CUDA already excluded.

Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants