Repository navigation
Conversation
Unreferenced leftover from the pre-dialog MCP design; references context props that no longer exist and breaks svelte-check. Assisted-by: pi
Drop the Router prefix from client-side API types; names now map directly to the /models endpoint family (load/unload/download/list). Merge ApiModelListResponse into ApiModelsListResponse (same endpoint shape in both modes) and remove the duplicate ModelsService.listRouter(). Assisted-by: pi
Add ModelDraftSidecar / ModelAuxSidecar enums with a ModelSidecar union type; mmproj is the only auxiliary sidecar (single member, covers vision and audio input). Add SIDECAR_PREFIX/SUFFIX_RE regex matching the server's filename conventions, and type guards + enum-file-token helpers in model-id.constants.ts. Extend parseModelId to detect sidecar filename tokens (mtp-, mmproj-, etc) and expose isDraftSidecar / isAuxSidecar / sidecarFromFileToken helpers. Add ModelCapability.TOOL_USE with icon/label/flag mappings. Assisted-by: pi
Use lowercase values for the sidecar enums so the value doubles as the filename token, derive the sidecar regexes from the enum values, and rename the MODEL_ID regex keys to the _REGEX suffix used by the rest of the constants files. Replace the tools capability magic string with ModelCapability.TOOL_USE. Assisted-by: pi
Add HuggingFaceService for browsing and searching GGUF models on the HF Hub: catalog/model search, model details, repo file tree, raw README fetch, and the llama.app model catalog. Includes GGUF file analysis helpers - extractQuantMeta (quant token plus sidecar type and its form, prefix or suffix), shard collapsing, quant bit-depth lookup, and download/size/likes formatting. Add the HF API types and the curated model list shown in the Discover Models sidebar. Assisted-by: pi
Assisted-by: pi
Address review on the HF data layer: - replace the HfModelSort / SidecarForm / sibling entry type string unions with enums (HfModelSort, SidecarForm, HfEntryType) - move URLs, query params, regexes, limits, retry settings, shard file conventions, tag tokens and formatting units into a dedicated huggingface.constants.ts; reuse the existing PATH_SEPARATOR - drop the task label / pipeline icon / library display maps: the discover UI only presents GGUF models, so keep the task tags for logic use only (parseTags) - drop the hardcoded curated model list; the discover dialog gets its default list from the llama.app /v1/catalog.json endpoint, which is an acceptable online-only source since the feature requires internet access anyway Assisted-by: pi
Address review follow-up: the llama.app catalog endpoint belongs to the models-discover feature, not the HF constants. Use Number() for shard index parsing and name the UD-quant prefix segment lookup. Assisted-by: pi
Port the hardware-compatibility estimator from ggml-org/llama-macos: map every GGUF file in a repo to a full/limited/none tier based on the device memory budget (GPU working set approximated from RAM, less fit slack and an OS floor) and the estimated weight + context memory. Main quants are tiered individually; shards, mmproj and quant-matched draft sidecars inherit their main quant's tier. Sidecar picking mirrors the server's find_best_sibling ranking (deepest directory, exact quant tag, closest bit depth). Also port detectToolUseSupport (infers tool-calling support from a chat template) and the browser get_info fallback helper. Assisted-by: pi
Replace the device-memory tier machinery with a plain file-size estimate: required runtime memory is the model file size with headroom for KV cache and allocator overhead (estimateModelMemoryBytes). Callers present the requirement; there is no device detection and no fit-versus-budget verdict. Drops resolveDeviceMemoryGb, deviceMemoryBudgetMb, computeFileCompatibilityTiers and the CompatibilityTier type, and the barrel keeps only the new estimator. Assisted-by: pi
Wire the model download flow: ModelsService.downloadModel (POST /models) and cancelDownload (DELETE /models), the apiDelete helper, ApiModelsDownloadRequest/Response types, the download_progress SSE payload, and the download_finished/download_failed SSE event kinds matching the server feed. Add modelsHubStore owning the HuggingFace GGUF model list for the discover dialog: curated catalog defaults on open, search replaces the list across all of HuggingFace. Assisted-by: pi
Port the download lifecycle into ModelsStatusManager: track per-entry progress keyed by <repo>:<tag> from the /models/sse feed, record failed downloads for the delete-and-retry path, and expose the downloadModel / cancelDownload operations (POST/DELETE /models). Add the ServerModelStatus.DOWNLOADED/DOWNLOADING cases and the ModelDownloadProgress type. Assisted-by: pi
Assisted-by: pi
Tag the pure-logic files and functions that llama.app (llama-pages) can reuse as-is: model id parsing, HF name and quant conventions, hardware compatibility estimation, chat-template capability detectors, and the HF formatting and metadata helpers. App-specific code is left unmarked. The LLAMA-APP-REUSE prefix makes the reusable surface greppable and distinguishable from regular comments: grep -rn LLAMA-APP-REUSE. Assisted-by: pi
Assisted-by: pi
Port the discover UI from the scrapbook, adapted to the typed sidecar API: searchable two-pane explorer (list, item, info, org avatar with quant badge), model details (header, name badges, download options grouped by bit depth with compatibility tiers, terminal serve/cli commands per draft sidecar, README viewer, chat template dialog), download confirmation dialog with progress, and the full-screen dialog shell. Presentational components take data and download state via props; the loading container and store wiring land in the integration branch. MarkdownContent gains a sanitized allowHtml option used by the README viewer; ModelId gains context, size-range, params and sidecar badges; the selector option passes thinking/tool flags instead of the removed capabilities prop. Basic Storybook stories cover each component with HF-shaped fixtures. Assisted-by: pi
Replace the device-memory tier badges with the simple memory estimate: each quant tooltip shows the estimated runtime memory and the device/OS chip is dropped, matching the estimateModelMemoryBytes model. Make the download dialog callbacks optional so the component stays presentational until wired to the live status store. Assisted-by: pi
Load the selected model's details, file tree and README via HuggingFaceService on selection change, and expose the download progress type globally for the status feed. The download options already read their state from the models status store. Assisted-by: pi
Add a Discover models item to the selector dropdown footer that opens the DialogModelsDiscover dialog. Assisted-by: pi
- ModelsDiscoverItem + ModelsDiscoverInfo fold into ModelsDiscoverListItem (avatar + model id + badges + context/size) - ModelsDiscoverDetails* renamed to ModelsDiscoverModelDetails* - TerminalCommands renamed to ModelsDiscoverModelDetailsCommands - ModelsDiscoverDetailsName folded into the details header - ModelsDiscoverListSearch extracted from the list search input - stories updated for the new names Assisted-by: pi
- ModelsDiscoverListSearch: extracted search input from the list - ModelsDiscoverModelDetailsMetadata: description + metadata chips extracted from the details header - ModelsDiscoverModelDetailsCommands: quant + draft sidecar selectors embedded in the inline command text - ModelsDownloadManager: tracked downloads with per-file progress and a delete action - ModelsDownloadManagerDownloadStatusToast: one toast per download with a progress bar per file (main + sidecars) and a CTA to open the download manager - DialogModelsDownloadManager: dialog shell for the manager Assisted-by: pi
The quant and draft sidecar buttons become a multiple toggle group (selecting which files to download, not download triggers), and below the panel the terminal llama serve command updates live to reflect the selection, with a download CTA that fires the downloads for the selected entries (main + optional draft sidecar). Adds the shadcn-svelte toggle-group component. Assisted-by: pi
Downloaded files are no longer selectable toggles; they render as a chip with a checkmark instead. Assisted-by: pi
extractQuantMeta parsed the full sibling path, so repo layouts that nest sidecars in a folder (e.g. MTP/mtp-Model-Q4_0.gguf) failed the sidecar prefix match and were classified as main weights. Assisted-by: pi:GLM-5.3-Flash
Replace the submenu-based dropdown with a flat scrollable list that has a sticky search header and actions footer. Add avatars and parameter badges to model options. Introduce ModelsSelectorReasoningPanel for in-place reasoning effort selection. Adjust dropdown sizing and input blur styling. Assisted-by: llama-ui:Qwen3.8-Flash-Next
Move the expand chevron to the right and swap it for ChevronUp when open. Group the effort label next to the title. Assisted-by: llama-ui:Qwen3.8-Flash-Next
The inline picks now mirror the selection one-way instead of seeding $state, the default quant seeds once when the file list resolves, and the command only renders while something is selected. Assisted-by: pi:zai-org/GLM-5.3
Assisted-by: pi:zai-org/GLM-5.3
Solo sidecar downloads register in /v1/models under the entry tag, so the state checks that first (normalized against the UD- quant prefix) and keeps the --model-draft / --mmproj args of registered models as a fallback. The chips no longer flip to downloaded while a download is in progress. Assisted-by: pi:zai-org/GLM-5.3
…cars A Q4_0-mtp style tag now resolves the sidecar file when no model file matches it, so a solo draft or mmproj download actually pulls the file. Cached sidecar files list as their own entries so the state survives a restart, and removing such a tag deletes only the sidecar. Assisted-by: pi:zai-org/GLM-5.3
Entries like repo:Q4_0-mtp mark a downloaded sidecar file, not a loadable model. Assisted-by: pi:zai-org/GLM-5.3
Replace the selector dropdown footer entry with a sidebar action that opens the discover dialog. Assisted-by: pi:zai-org/GLM-5.3
Move the option's fallback logic into useModelParamsFallback and reuse it for the dropdown and sheet trigger badges. Assisted-by: pi:zai-org/GLM-5.3
Hide the ModelId badge and icon wrappers when empty, normalize its indentation, and tighten the reasoning panel and discover detail spacing. Assisted-by: pi:zai-org/GLM-5.3
Every quant chip is now an independent download action with its own lifecycle: download, retry, pause, resume, cancel and delete with confirmation. The store distinguishes user-stopped downloads over the /models/sse feed so they settle silently, keeps paused progress resumable, resolves the server-registered id on delete and refetches the model list so the selector stays in sync with downloads. Drops the selection toggle group, the download CTA and the download manager dialogs and progress toasts. The serve command preview is now a standalone widget with its own picks and an addable draft segment. Assisted-by: pi
A "Download in progress" section above the loaded models lists the in-flight and paused downloads with a live progress bar and the same pause, resume and cancel actions as the discover quant chips, with the same avatar and model id presentation as the regular option rows. Assisted-by: pi
Assisted-by: pi
The download monitor thread acquires the mutex on its way out, so joining it while holding the lock in server_models::remove deadlocks once the status has flipped to DOWNLOADED. Join outside the lock, same pattern as load_models(). Assisted-by: pi:zai-org/GLM-5.3
extractQuantMeta now recognizes a -draft tail after the sidecar token (Model-MTP-draft.gguf), a bare sidecar token filename (imatrix.gguf) and suffix sidecars whose head carries no quant (Model-imatrix.gguf). imatrix is an aux sidecar: it stays in the download chips with a badge, is never a serve-command option, and mmproj alone drives the Vision capability. Assisted-by: pi:zai-org/GLM-5.3
Cancelling stops the download and discards the partial files, so the selector item and the download chip both ask for confirmation first. Assisted-by: pi:zai-org/GLM-5.3
Favorites sit above the in-flight downloads; the dropdown visual order follows the new list order. Assisted-by: pi:zai-org/GLM-5.3
Downloading entries are not usable models yet; the selector tracks them in its download-progress section instead. Assisted-by: pi:zai-org/GLM-5.3
Assisted-by: pi:zai-org/GLM-5.3
The copy button swaps to a checkmark briefly; the close X moves into the header row so it lines up with the title. Assisted-by: pi:zai-org/GLM-5.3
Move the ModelsDiscover components into ModelsDiscoverList/ and ModelsDiscoverDetails/ subfolders (with ModelsDiscoverDetailsDownloadOptions/ nested inside), and the model selector components into a ModelsSelector/ subfolder. Rename ModelsDiscoverModelDetails* to ModelsDiscoverDetails* for consistency, add per-folder index.ts barrels, and keep shared leaves (ModelId, ModelBadge, ModelLoadHighlight, ModelsDiscoverAvatar, DownloadProgressBar) at their folder roots.
allozaur
left a comment
There was a problem hiding this comment.
Review notes — memory-fit UI info
Goal for this area: keep only the same memory-fit UI logic as the llama.app models pages, and introduce no device-memory checking in our changes.
Good news — no device-memory logic remains. The device-tier machinery (resolveDeviceMemoryGb, deviceMemoryBudgetMb, computeFileCompatibilityTiers, CompatibilityTier, navigator.deviceMemory, the full/limited/none verdict) was dropped in #27957 (a07359ef) and never came back on the stack base.
One thing to reconcile — tools/ui/src/lib/utils/model-compatibility.ts currently exposes two functions:
estimateModelMemoryBytes()— the simplified size × 1.05; appears unused in the discover UI now.minMemoryTierGb()— re-added on this branch (46c1d39b1); functionally identical to llama.app'sModelsCatalogService.minMemForBuild(sameMAC_MEM_TIERS,RAM_BUDGET_RATIO = 0.75,RAM_OVERHEAD_MB = 2048,QUANT_WEIGHT = 1.05). It is a pure file-size → required-tier function (no device detection), rendered as "needs at least NGB+ memory" in the download-options row.
minMemoryTierGb is the keeper — it matches the llama.app behavior. Suggest dropping estimateModelMemoryBytes if it's truly unused, so we expose a single memory-fit function that mirrors minMemForBuild.
No action needed on device logic — it's already clean.
d448211 to
1884398
Compare
Overview
WIP layer on top of
allozaur/ui-reuse-markersfor the Discover & Download Models feature.Models Discover Hub
Download lifecycle
llama servecommand preview with inline selects and copy<quant>-<sidecar>tag resolution, cached-sidecar listing/removal, deadlock fixComponent structure
ModelsDiscoverList/,ModelsDiscoverDetails/(with nestedModelsDiscoverDetailsDownloadOptions/) andModelsSelector/folders with per-folder barrels