Repository navigation
Add Clef as a local System One referee, and pin llama.cpp b11485 - #4
Draft
markelphoenix wants to merge 14 commits into
Draft
markelphoenix wants to merge 14 commits into
markelphoenix wants to merge 14 commits into
Conversation
Clef and Clef-flash are offered after the story model is ready, with published GGUF sizes and an Apache-2.0 notice. The pinned llama.cpp build cannot load them, so the menu explains that and does not download the weights. Hardware matching now names multi-GPU fits honestly, and the download confirm screen states the hardware, warranty, and output limits. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
The pin is the newest release that still publishes the five archives the game builds use. An older installed engine is offered that download instead of stopping. Each graphics card keeps its own memory reserve, multi-GPU speed is a layer split, and a fast card-resident referee is not passed over for a much slower CPU fit. The AI-output line stays objective: output can be wrong, and the player is responsible for how they use it. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
TERM=dumb was forcing 80 columns even when a width was set, and NO_COLOR was stripping colour the tests asked for. The harness treats that terminal as capable. The smoke script, which starts its own Python, gets wide columns and no colour so its engine line stays one sentence. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
b11485 turns --fit on with a 1024 MiB device margin, which would spill layers the menu said still fit. GPU launches now pass the menu's 819 MiB reserve and the planned context. CPU launches pass --fit off, because the fitter treats leftover RAM as unlimited. Clef's head reads rows of a quantized output tensor. The default repack buffer cannot do that and the server aborts before it listens. Embedding mode also shrinks the batch to 512, which is smaller than a real referee prompt. Local Clef now passes --no-repack and a physical batch equal to the context. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
Tests stub NVML, the registry, and DXGI so they cannot load a real GPU. Clef's physical batch stays at most 4,096 while the context stays large, and a request that would not fit is trimmed before it is sent. The client waits longer when the model is off the GPU, and a local timeout talks about hardware speed. The fit budget and --fit-target subtract memory already in use, with extra margin on Windows. A tighter machine no longer recommends the larger referee, and the menu shows seconds per decision. The engine pin covers every archive the Clef installer can download, and a checksum mismatch fails closed. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
A 3 GB cut on every Windows GPU made a 6 GB card too small for the 4B story model the store page promises. Story budgets still subtract memory already in use. Clef's budget and --fit-target keep the extra 3 GB on Windows, which is where the measured run overran the projection. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
Discovery now keeps image-text-to-text and untagged GGUF repos, still skipping mmproj files and every existing license and publisher check. unsloth/Qwen3.8-27B-GGUF is a seed so a 32 GB card has an official 27B option; a comfortable Qwen3 32B stays the pick. LoRA paths are not treated as full weights. When the file size and parameter count are known, quality and fit use bytes times 8 divided by parameters, and an unknown label scores low until the size is known. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
Measured bits still score an unknown or last-resort name, and a file that is really more compressed than its label. A Q4_K_M that only measures high because of the tokenizer stays on the 4-bit rung, so an 8 GB laptop still gets Clef-flash and a Mac still prefers the quant that fits on the GPU. The built-game architecture list includes qwen35. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
A laptop's built-in chip is skipped even when it reports more memory than the discrete card. Steam Deck's Custom GPU 0405 is shared RAM, so the plan stays on the CPU. A fast split beats a slow CPU story, and a partial Clef that is several times quicker replaces a CPU referee. The survey locks a story and a referee for the common Steam tiers. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
…unch. When in-use video memory was never read, keep about 2 GiB spare. Class a corrupt model file ahead of a GPU switch, and name the CPU instruction the probe can see is missing. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
The launch check reads the version word after the magic. A zero version is refused, so the fake model now says version 3. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
The story launch already passes --device. Clef did not, so --fit could spread onto the story card or a built-in chip. Reuse that build's device listing when the engine file matches, or read --list-devices once, and pin the referee to the cards left after the story model is reserved. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
RTX Spark, DGX Spark, and GB300 use a share of system RAM, not VRAM plus RAM, with an LPDDR5X-class speed guess. Apple Silicon keeps a recommended-working-set share and Metal. Strix Halo and Lunar Lake use dedicated plus shared memory, capped at that share. Windows and Linux Arm stay on the pinned arm64 CUDA, Vulkan, or CPU build. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
The Clef launch checks the GGUF header before the process starts. On a 5090 with the 32B story loaded and about 2.1 GiB already in use, both referees are on the CPU, and the faster file wins. Linux --fit-target no longer adds in-use memory that CUDA already excluded. Co-authored-by: markelphoenix <markelphoenix@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is not legal advice. The notices in this pull request are plain-language warnings. They do not create a complete EULA, and they should not be treated as sufficient legal protection. A lawyer should review the items listed under "Remains for human or legal review" before this is relied on for a store page or a public release.
Do not merge from this draft.
Screenshots below are from an earlier round. They still show a tokens/s figure. The menu now shows estimated seconds per decision. This environment has no Tk, so they are not a screenshot of the game window. The pin in these pictures is llama.cpp b11485.
System One on a 32 GB RTX 5090, previous round. The speed line now says seconds per decision.
System One on an 8 GB laptop GPU. Clef-flash Q4_K_M is recommended. Full Clef is a snug fit.
Installed engine b11100 is too old for Clef. The menu offers to download llama.cpp b11485.
Steam Hardware Survey
The fit is checked on simulated machines from the common Steam range, not only the RTX 5090 test PC. Each row locks the story model and the System One referee. Download size plus 1 GB of spare disk stays inside the free disk in the fixture (200 GB). Estimated memory stays inside the planner's budget. That budget already keeps 0.8 GB of video memory spare, subtracts memory other programs are using, and keeps an extra 3 GB spare for Clef on Windows only. Story models do not take that extra cut.
A recommended story is at least 3 tokens/s for a whole turn and is not labelled "very slow". The speed word matches the raw tokens/s estimate. A local referee's estimated seconds stay under the timeout for where it runs: 30 s on the GPU or in one shared-memory pool, 120 s for a GPU+RAM split, 180 s on the CPU.
"Memory" is need / budget. On a discrete GPU the budget is usable video memory. On a unified-memory machine it is a share of system RAM (65% at 8 GB, 70% through 48 GB, 75% at 64 GB, 80% at 128 GB), and that share is not added to RAM again. On the CPU, and for the RAM side of a split, it is spare system RAM. A split also says what share of the layers stays on the card.
Windows 10 and Windows 11 use the same fit rules. Both are in the matrix so a version string cannot silently change the pick. AMD and Intel Arc rows are the Vulkan path: no NVML and no nvidia-smi. Apple Silicon is included because the game ships a Metal build. The later rows add more Mac sizes, RTX Spark and DGX Spark, Ryzen AI Max, and a Lunar Lake laptop. Those budgets are a share of system RAM.
The rows that were already in this table are unchanged by the unknown-memory margin. Each survey card states its in-use figure, including zero, so the 0.8 GB reserve still applies. A probe that never read in-use memory is a separate case and keeps about 2 GB spare.
Picks that were wrong, and what they are now
Expected picks
When local AI does not fit
mock).The two 12 GB NVIDIA rows (Windows 11 and Linux) pick the same story and the same referee. Linux keeps 2.5 GB of RAM headroom and Windows keeps 3.5 GB, so the RAM budget differs (29.5 GB vs 28.5 GB) and the pick does not. The dual 3060 pools both cards: usable video memory is
2 × (12 − 0.8)GB.Unified memory
Apple Silicon already planned inside one pool. This round uses that for the other machines whose CPU and GPU share RAM, and it tightens the Apple share on small and large Macs.
The GPU budget is a share of system RAM: 65% at 8 GB, 70% through 48 GB, 75% at 64 GB, and 80% at 128 GB. That percentage is the headroom. The fit does not also subtract the OS reserve from the share, and it does not add the share to RAM a second time. A BIOS "dedicated VRAM" number is not the budget by itself. When DXGI reports dedicated plus shared, the budget is the sum, capped at the share. A carve-out with no shared figure uses the share and ignores the carve-out. nvidia-smi reporting N/A, or reporting the whole RAM stick, does not become a second pool. Used memory against a whole-pool reading is scaled down to the share.
Steam Deck's Custom GPU 0405 stays shared RAM with a CPU plan. A discrete card still wins over a built-in chip, including a named NVIDIA card whose size was never read. A lone Iris Xe, UHD, Phoenix, or 780M detected live is now that RAM share (Vulkan when the engine can use it). The Iris Xe and 780M rows already in the table still pass zero video memory, so those CPU picks stay locked.
Speed for RTX Spark, DGX Spark (GB10), and GB300 uses an LPDDR5X-9400 class guess, 301 GB/s, not a discrete-GPU table. Apple chips keep their M1–M4 Pro/Max/Ultra bandwidth tiers, and the engine is Metal. Strix Halo (Radeon 8060S) stays at 256 GB/s. Lunar Lake Arc 140V stays at 136 GB/s. Strix Halo and Lunar Lake prefer Vulkan.
Windows on Arm with an NVIDIA GPU plans CUDA 13 arm64 when the driver is 580 or newer (or unknown), then Vulkan arm64, then CPU arm64. An older driver starts at Vulkan arm64. Linux arm64 with NVIDIA is CUDA 13 arm64, then Vulkan arm64 when the loader is present, then CPU arm64. The setup menu names that build and says an x64 engine is not run under emulation.
select_assetsmatches the host architecture, so an arm64 plan cannot return an x64 archive name.b11485 publishes those arm64 archives. The pin already had their SHA-256 lines, and this round checked them against the release API. They match. There is no CUDA 12 arm64 build.
llama-b11485-bin-win-cuda-13.4-arm64.zipandcudart-llama-bin-win-cuda-13.4-arm64.zipllama-b11485-bin-ubuntu-cuda-13.4-arm64.tar.gzandcudart-llama-b11485-bin-ubuntu-cuda-13.4-arm64.tar.gzllama-b11485-bin-win-vulkan-arm64.zipandllama-b11485-bin-ubuntu-vulkan-arm64.tar.gzllama-b11485-bin-win-cpu-arm64.zipandllama-b11485-bin-ubuntu-arm64.tar.gzThe survey rows use the share (22.4 GB on 32 GB, 48 GB on 64 GB, 102.4 GB on 128 GB). Two detection cases are tighter than that share. A 128 GB Strix Halo whose DXGI numbers are 64 GB dedicated and 32 GB shared is budgeted at 96 GB. A 32 GB Lunar Lake whose DXGI numbers are 0.125 GB dedicated and 15.5 GB shared is budgeted at 15.6 GB. The matrix row for Lunar Lake is the 22.4 GB share, which is what the fit uses when Windows does not report a shared figure beside a tiny carve-out.
GB300 uses the same budget and the same 301 GB/s guess as DGX Spark. The survey table does not add a separate GB300 row.
Model discovery (already on this branch)
Three model-discovery bugs, read against commit d62ef5f. The six hardware-test fixes from the previous round stay as they are. Abliterated and uncensored repos are still rejected. The family-friendly filter and the Steam disclosure are unchanged. Underdog Saluki is not a catalog entry; its byte and parameter counts are only a test of the bits-per-weight math. The GSQ-RCO repo is not a catalog entry either.
Qwen3.8 27B GGUF repos can be found. Hub search used
filter="gguf"withpipeline_tag="text-generation". The major Qwen3.8-27B GGUF repos are taggedimage-text-to-textor have no pipeline tag (unsloth,ggml-org,bartowski,lmstudio-community). Each publisher search now also asks forimage-text-to-textand for repos with no pipeline tag, then keeps text-generation, image-text-to-text, and untagged results. They can still be used for text only.mmprojfiles are still skipped. License, publisher, and specialist checks are the same, so a vision-named, speech, coder, or abliterated repo is still dropped.unsloth/Qwen3.8-27B-GGUFis a seed: Apache-2.0, baseQwen/Qwen3.8-27B, architectureqwen35(llama.cpp b11485 loads it), 27.32B parameters, KV shape 64 layers × 4 KV heads × 256. The published files are Unsloth's UD variants; the seed stores those sizes under the plain quant names (Q8_029.05 GB,Q6_K21.98,Q5_K_M19.77,Q4_K_M16.46,IQ4_XS14.25). There is noQ3_K_Mfile in that repo, so the seed does not invent one. A built game that can load the curated list treatsqwen35as known, so the seed is offered.On a 32 GB card with 64 GB RAM the recommendation stays Qwen3 32B (on the card). Qwen3.8 27B at Q5_K_M is close behind, also on the card, with a lower score. On a 12 GB card both spill, the 27B spill scores higher than the 32B spill, and the pick is Qwen3 14B, which still fits on the card.
A LoRA adapter is not a full model.
parse_quantandpick_gguf_fileskip any path containinglora(case-insensitive), next to the existingmmprojskip. A path likelora/NAME-LoRA-r320-F16.ggufis no longer returned as an F16 weight.Quality follows the file, not a misread label.
parse_quantacceptsIQn-mixandIQn_mix.entry_from_hubkeeps that file. When the file size and the parameter count are known, bits per weight arefile bytes × 8 ÷ parameter count. The label is used only when the size is unknown. An unknown label scores 0.55, not the old 0.9. Saluki's published figures (7,898,369,152 bytes, 26,895,998,464 parameters) are about 2.35 bits per weight and stay a last resort. ISTA-DASLab's GSQ-RCO file namedIQ3_S(11,771,546,784 bytes, same parameter count, about 3.50 bits) is the low rung, so a comfortable fit is not forced into "heavily compressed."A known quant is not scored above its name when the file is only a little larger than the label (tokenizer and metadata). Clef's Q4_K_M stays on the 4-bit rung, so an 8 GB laptop still gets Clef-flash, and a Mac still prefers the quant that fits on the GPU. A last-resort name whose file is really at least 3.5 bits is promoted. A file that measures worse than its name uses the measured score, so a 2.35-bit file cannot hide behind a Q4 label.
Hardware round (already on this branch)
Lux's Windows 11 machine (Ryzen 7 9850X3D, 64 GB RAM, RTX 5090 32 GB, driver 617.14, Clef Q4_K_M, llama.cpp b11485 CUDA 13) scored 5/5 prompts at context 4096 in 0.35–0.88 s and used 19.1 GiB. Context 16384 also worked. Six bugs from that run are fixed here.
Tests no longer touch a real GPU. The specs fixture and an autouse guard stub NVML, the Windows registry, and DXGI. If a test still reaches the real probe, it raises before any DLL loads. On that PC, four specs tests failed because the NVML fallback loaded
nvml.dll. Those tests pass with the stubs. On that round,tests/test_specs.pywas 60/60 on this Linux VM. The survey adds the Steam Deck case on top of that.The physical batch is capped at 4,096.
-ccan stay large.-band-ubaremin(context, 4096). Matching them to a 65,536 context made the compute buffer 26,519 MiB and CUDA failed. The same 65,536 context with a 4,096 batch stayed on the GPU at about 18.7 GB. The request builder trims a prompt that would exceed the batch, so the server is not asked to score something the batch cannot hold. Clef's context cache measured 0 MiB, so the planner no longer grows a KV estimate with-c. The compute buffer in the estimate stays about 1.3 GB, the size measured at batch 4,096.The referee waits longer off the GPU. A full-GPU call still times out at 30 s. A GPU+RAM split waits 120 s. A CPU call waits 180 s. The measured CPU rounds were 47–65 s, so every one of them used to fall back to the local judge. A local timeout says the hardware is busy. The hosted Jev message is unchanged.
Free memory, not the card's total. nvidia-smi / NVML in-use memory is subtracted from every menu fit budget. Windows adds that same in-use figure to
--fit-target, because CUDA there had reported about 30,991 MiB free while other programs were using 1–6 GB. Linux--fit-targetstays the 0.8 GiB reserve: CUDA's free figure on Linux already excludes other programs, so adding them again would keep that memory spare twice. "Needs ~X of Y GB" uses the free figure. An extra 3 GB on Windows applies to Clef's budget and Clef's--fit-targetonly, because that run used about 2.9 GB more than llama.cpp projected. Story models do not take that cut: a 6 GB Windows card would otherwise miss the 4B story model the store page promises. A Windows story launch still adds in-use memory to--fit-target. With nothing in use and no specs, the documented target stays 819 MiB.More memory pressure does not pick a bigger referee. A GPU+RAM split is a worse home than a model that still fits wholly on the card. With Qwen3 8B loaded, full Clef was a snug full-GPU fit and Clef-flash won. With Qwen3 14B loaded, full Clef became a 75% split and used to win instead. The larger model no longer wins that comparison. The menu shows estimated seconds per decision (about 0.6 s fully on the card, about 1 s at a 75% split, about a minute on 8 CPU cores for Clef Q4), not tokens/s. A real 75% split (
-ngl 48) took about 1.0 s per round. The story menu's tokens/s column is unchanged.The engine pin covers every archive the Clef installer can download. SHA-256 lines were added for the CUDA 13 and CUDA 12 builds, the CUDA runtime zips, and the Vulkan and CPU archives for each OS that path can select. Windows CUDA runtime zips omit the build tag in the filename; the pin parser accepts that.
install_tagged_releasechecks the pin before download and fails closed when the pin has no digest or GitHub's digest differs. The five game-build digests are unchanged. The new lines are thesha256values GitHub publishes on the b11485 release. Those zip files were not downloaded again on this machine.(a) What changed, including earlier rounds
Clef is Cloudflare's open-weight decision model. It scores labelled answers in one forward pass and does not write free-form story text, so it is a System One referee, not a storyteller. It is not in
MODEL_CATALOG.SYSTEM_ONE_CATALOG: Clef (27.02B,ggml-org/Clef-GGUF) and Clef-flash (9.08B,ggml-org/Clef-Flash-GGUF). Quant sizes are the published GGUF byte lengths. License is Apache-2.0. Requested context is 4,096 tokens. Published native context is 65,536. Architecture isclef.recommend_system_oneprefers the larger comfortable model on the best home. A split does not count as "on the card." A card-resident fit at least four times quicker, in seconds per decision, replaces a larger CPU fit. A partial split counts as card-resident for that comparison.--no-repack,--fit onwith the margin above, and-b/-ubcapped at 4,096. CPU launches pass--fit off,--device none, and-ngl 0. A GPU or split launch also passes--devicefor the cards left after the story model is reserved.unsloth/Qwen3.8-27B-GGUF, Apache-2.0, architectureqwen35). Discovery can also see image-text-to-text and untagged GGUF repos.(b) Evaluation notes that this round changes
lora/...-F16.ggufwas chosen as a full F16 weight.lora.IQ2-mixwas dropped, and a label could overrate a 2-bit file or underrate a 3.5-bit file namedIQ3_S.IQn-mix. Use measured bits per weight when size and parameter count are known. Unknown labels score 0.55. A known name is not promoted by tokenizer overhead.nvml.dll.-b/-ubequal to a large context filled the card.--fit-target 819ignored memory already in use, and the projection was short on Windows.--fit-targetadds it; Linux stays at 819 MiB. Add 3 GB for Clef on Windows only.Story-catalog spot checks that already looked right were left alone. The story recommender's 8 tokens/s floor for the first tier is still intentional. A remembered Clef choice still asks again instead of auto-starting. LEARN's hosted-Jev pages still say Jev.
(c) Liability checklist
This is not legal advice.
Found, then kept
Added or reworded
mmprojfiles are still skipped. ARCHITECTURE also says the primary card is the discrete GPU, and that Steam Deck's Custom GPU 0405 shares its memory with the CPU.Remains for human or legal review
(d) Sources
qwen35)Workers AI was not called. No third-party VRAM calculator was used. The new pin lines were copied from GitHub's published digests, not hashed again from the zip bytes on this machine. The arm64 CUDA 13.4, Vulkan, and CPU archives for Windows and Ubuntu were already in that pin. This round checked those digests against the b11485 release API again. They match. There is no CUDA 12 arm64 archive. Those zip bytes were not downloaded or run. The original five game-build archives were hashed from the downloaded bytes in an earlier round and already matched those digests. Qwen3.8 file sizes and the GGUF parameter count were read from the Hub listing. Saluki and the GSQ-RCO
IQ3_Sfile were used only as test figures; neither repo was added. The survey machines are simulated specs. They were not measured on those cards.Clef on the earlier CPU smoke
Clef-flash Q4_K_M did run here, on the Ubuntu CPU build, before the hardware round. Three geese prompts scored in 54–60 s with
output_tokens: 0. That smoke used-b/-ub2048 at context 2048. The game now caps the batch at 4,096 instead of matching whatever context was requested.Planned card, unknown VRAM, and a bad file
These checks were reimplemented in this repository. No code, patterns, or tests were copied from the AGPL-3.0 project that suggested them.
--list-devicesis matched to the cards the fit already pooled (the primary card, plus other cards from that vendor with at least 4 GB). The launch passes--devicefor those rows. Two NVIDIA cards stayCUDA0,CUDA1, withCUDA_DEVICE_ORDER=PCI_BUS_IDandCUDA_VISIBLE_DEVICES=0,1. A Vulkan listing that also shows an Intel or AMD built-in chip passes only the discrete card. A loneCUDA1is renamed toCUDA0once it is the only visible device. The list-devices probe sets the PCI order for NVIDIA and does not setCUDA_VISIBLE_DEVICES. The Clef process uses that listing when the engine file is the same, or reads--list-devicesonce for its own build. It is pinned to the cards left after the story model's video memory is reserved. That can be a different card: a Vulkan build that sees an RTX 4090 and an RX 7800 XT starts the story on the 4090 and, once that card is reserved, starts Clef on the 7800 XT. Two identical cards keep their list position, so a referee planned on the second card isCUDA_VISIBLE_DEVICES=1(the engine sees it asCUDA0). A built-in chip in the same listing is still left out.vram_used_knownis false when nvidia-smi had no used column, NVML returned no used figure, Linux AMD had nomem_info_vram_used, or the size came from the Windows registry or DXGI. The menu fit then keeps 2 GiB instead of 0.8 GiB plus a measured in-use value. On Windows,--fit-targetdoes the same, and the Clef extra 3 GB is still added on top. On Linux,--fit-targetstays 819 MiB, because there is no separate CUDA-free reading and the free figure already excludes other programs. Linux AMD readsmem_info_vram_usedand clamps it to the total.GGUFand a version from 1 to 3. The message tells the player to delete it and download it again, and says the engine was not started. The same check runs for anmmprojpath when one is set. This game does not download a projector. Clef's launch runs that check after the download returns and before the process starts, and raisesLocalClefUnavailablewith the same refusal.gguf_init_from_(any reader),wrong number of tensors, andfailed to open ggufare classified as the model file before a CUDA or Vulkan error line. A CUDA out-of-memory line that also saysfailed to load modelstays a GPU error. Those markers match the public llama.cpp phrases. This machine did not replay a corrupt file on b11485./proc/cpuinfonow recordsbmi2). An empty flag list or a NEON-only chip keeps the generic sentence. Windows has no standard check for FMA, F16C, or BMI2, so those are not reported as missing there.Nothing in those seven items was already implemented. The survey table did not move for that round.
ALUCARD rerun
Lux reran the game at 9a092bc on the Windows RTX 5090 (ALUCARD). That run passed: Windows pytest 3779 passed and 0 failed; Clef at 4096, 16384, and 65536 context all launched with
-b/-ub4096, peaked at 21,746 MiB, and scored in 0.3–0.8 s; the device pin gave CUDA0 and left out the AMD iGPU; the GGUF check works; checksums fail closed; and with VRAM in use the planner picked Clef-flash at 86% on the GPU, with both models peaking at 27.7 of 32.6 GB. Three gaps from it are fixed here. The unified-memory commit did not move the 5090 survey row: with nothing in use, the story is still Qwen3 32B Q5 on the GPU and the referee is still Clef-flash Q4 with about 59% of its layers on the card (about 2.2 s / 120 s).download_ggufreturns, thengguf_file_problemruns, and a bad file raisesLocalClefUnavailablewith the same refusal the story engine uses. The process is not started.--fit-targetis 6042 MiB. Qwen3 32B Q5 still sits fully on the GPU at about 23.5 GB, and both Clef files then only fit on the CPU (about 55 s and about 19 s on 8 cores). A partial-GPU Clef-flash still beats a CPU referee when the card can hold one. When both are on the CPU, the faster file wins if the larger is more than twice as slow or over about 30 s. The same rule moves the other CPU-only Clef rows in the table to Clef-flash.--fit-targetdoes not add in-use memory. Windows CUDA's free figure ignores other programs, so that launch still adds them. Linux CUDA's free figure already excludes them. The menu budget still subtracts in-use memory on both. There is no stored CUDA-free field, so the split is the OS.Tests
TERM=dumb PYTHONPATH=src python3 -m pytest: 3768 passed, 60 skipped, 0 failed (33.11s). The previous full suite was 3766 passed, 60 skipped. The 5090 survey row is unchanged: Clef-flash Q4 with about 59% of its layers on the card, about 2.2 s / 120 s. The unified-memory commit did not move that row. Twelve rows that had full Clef on the CPU now pick Clef-flash on the CPU, because that file is more than twice as fast. One earlier run of the RAM-bandwidth buffer test hit the known noisy-neighbour ratio (about 2.05 against a limit of 2.0) and was not loosened.Not verified on this machine
vram_used_gbon the simulated card.--fit-target, including the extra 3 GB on a Clef launch.cleforqwen35.unsloth/Qwen3.8-27B-GGUF. The three-query search is covered by a fake Hub client.--deviceandCUDA_VISIBLE_DEVICESon a real dual-GPU or Vulkan machine, including the Clef process. The story and referee cases are simulated listings.mem_info_vram_usedon a real AMD GPU.recommendedMaxWorkingSetSize. The shares and the 301 GB/s LPDDR5X-9400 guess are simulated.To show artifacts inline, enable in settings.