Skip to content

ui : Models Discover and Download (WIP) - #28374

Closed
allozaur wants to merge 65 commits into
ggml-org:allozaur/ui-reuse-markersfrom
allozaur:allozaur/discover-wip
Closed

allozaur wants to merge 65 commits into
ggml-org:allozaur/ui-reuse-markersfrom
allozaur:allozaur/discover-wip

Conversation

@allozaur

@allozaur allozaur commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Overview

WIP layer on top of allozaur/ui-reuse-markers for the Discover & Download Models feature.

Models Discover Hub

  • Sidebar "Discover models" entry opening a two-pane dialog (searchable list + details)
  • Curated llama.app catalog by default; HuggingFace full-text GGUF search
  • Detail pane: header, metadata chips, chat template dialog, README

Download lifecycle

  • Per-quant download chips (download/pause/resume/cancel + progress)
  • llama serve command preview with inline selects and copy
  • In-flight/paused downloads in the model selector
  • Server: <quant>-<sidecar> tag resolution, cached-sidecar listing/removal, deadlock fix

Component structure

  • Reorganized into ModelsDiscoverList/, ModelsDiscoverDetails/ (with nested ModelsDiscoverDetailsDownloadOptions/) and ModelsSelector/ folders with per-folder barrels

Draft: used for code review; not yet split into dedicated stack branches.

Unreferenced leftover from the pre-dialog MCP design; references
context props that no longer exist and breaks svelte-check.

Assisted-by: pi
Drop the Router prefix from client-side API types; names now map
directly to the /models endpoint family (load/unload/download/list).
Merge ApiModelListResponse into ApiModelsListResponse (same endpoint
shape in both modes) and remove the duplicate ModelsService.listRouter().

Assisted-by: pi
Add ModelDraftSidecar / ModelAuxSidecar enums with a ModelSidecar
union type; mmproj is the only auxiliary sidecar (single member,
covers vision and audio input). Add SIDECAR_PREFIX/SUFFIX_RE regex
matching the server's filename conventions, and type guards +
enum-file-token helpers in model-id.constants.ts.

Extend parseModelId to detect sidecar filename tokens (mtp-, mmproj-,
etc) and expose isDraftSidecar / isAuxSidecar / sidecarFromFileToken
helpers. Add ModelCapability.TOOL_USE with icon/label/flag mappings.

Assisted-by: pi
Use lowercase values for the sidecar enums so the value doubles as
the filename token, derive the sidecar regexes from the enum values,
and rename the MODEL_ID regex keys to the _REGEX suffix used by the
rest of the constants files. Replace the tools capability magic
string with ModelCapability.TOOL_USE.

Assisted-by: pi
Add HuggingFaceService for browsing and searching GGUF models on the
HF Hub: catalog/model search, model details, repo file tree, raw
README fetch, and the llama.app model catalog. Includes GGUF file
analysis helpers - extractQuantMeta (quant token plus sidecar type
and its form, prefix or suffix), shard collapsing, quant bit-depth
lookup, and download/size/likes formatting.

Add the HF API types and the curated model list shown in the
Discover Models sidebar.

Assisted-by: pi
Address review on the HF data layer:

- replace the HfModelSort / SidecarForm / sibling entry type string
  unions with enums (HfModelSort, SidecarForm, HfEntryType)
- move URLs, query params, regexes, limits, retry settings, shard
  file conventions, tag tokens and formatting units into a dedicated
  huggingface.constants.ts; reuse the existing PATH_SEPARATOR
- drop the task label / pipeline icon / library display maps: the
  discover UI only presents GGUF models, so keep the task tags for
  logic use only (parseTags)
- drop the hardcoded curated model list; the discover dialog gets its
  default list from the llama.app /v1/catalog.json endpoint, which is
  an acceptable online-only source since the feature requires internet
  access anyway

Assisted-by: pi
Address review follow-up: the llama.app catalog endpoint belongs to
the models-discover feature, not the HF constants. Use Number() for
shard index parsing and name the UD-quant prefix segment lookup.

Assisted-by: pi
Port the hardware-compatibility estimator from ggml-org/llama-macos:
map every GGUF file in a repo to a full/limited/none tier based on
the device memory budget (GPU working set approximated from RAM, less
fit slack and an OS floor) and the estimated weight + context memory.

Main quants are tiered individually; shards, mmproj and quant-matched
draft sidecars inherit their main quant's tier. Sidecar picking
mirrors the server's find_best_sibling ranking (deepest directory,
exact quant tag, closest bit depth).

Also port detectToolUseSupport (infers tool-calling support from a
chat template) and the browser get_info fallback helper.

Assisted-by: pi
Replace the device-memory tier machinery with a plain file-size
estimate: required runtime memory is the model file size with
headroom for KV cache and allocator overhead (estimateModelMemoryBytes).
Callers present the requirement; there is no device detection and no
fit-versus-budget verdict.

Drops resolveDeviceMemoryGb, deviceMemoryBudgetMb,
computeFileCompatibilityTiers and the CompatibilityTier type, and the
barrel keeps only the new estimator.

Assisted-by: pi
Wire the model download flow: ModelsService.downloadModel (POST
/models) and cancelDownload (DELETE /models), the apiDelete helper,
ApiModelsDownloadRequest/Response types, the download_progress SSE
payload, and the download_finished/download_failed SSE event kinds
matching the server feed.

Add modelsHubStore owning the HuggingFace GGUF model list for the
discover dialog: curated catalog defaults on open, search replaces
the list across all of HuggingFace.

Assisted-by: pi
Port the download lifecycle into ModelsStatusManager: track per-entry
progress keyed by <repo>:<tag> from the /models/sse feed, record
failed downloads for the delete-and-retry path, and expose the
downloadModel / cancelDownload operations (POST/DELETE /models).

Add the ServerModelStatus.DOWNLOADED/DOWNLOADING cases and the
ModelDownloadProgress type.

Assisted-by: pi
Tag the pure-logic files and functions that llama.app (llama-pages)
can reuse as-is: model id parsing, HF name and quant conventions,
hardware compatibility estimation, chat-template capability
detectors, and the HF formatting and metadata helpers. App-specific
code is left unmarked.

The LLAMA-APP-REUSE prefix makes the reusable surface greppable and
distinguishable from regular comments: grep -rn LLAMA-APP-REUSE.

Assisted-by: pi
Port the discover UI from the scrapbook, adapted to the typed
sidecar API: searchable two-pane explorer (list, item, info, org
avatar with quant badge), model details (header, name badges,
download options grouped by bit depth with compatibility tiers,
terminal serve/cli commands per draft sidecar, README viewer, chat
template dialog), download confirmation dialog with progress, and
the full-screen dialog shell.

Presentational components take data and download state via props;
the loading container and store wiring land in the integration
branch. MarkdownContent gains a sanitized allowHtml option used by
the README viewer; ModelId gains context, size-range, params and
sidecar badges; the selector option passes thinking/tool flags
instead of the removed capabilities prop.

Basic Storybook stories cover each component with HF-shaped
fixtures.

Assisted-by: pi
Replace the device-memory tier badges with the simple memory estimate:
each quant tooltip shows the estimated runtime memory and the device/OS
chip is dropped, matching the estimateModelMemoryBytes model. Make the
download dialog callbacks optional so the component stays presentational
until wired to the live status store.

Assisted-by: pi
Load the selected model's details, file tree and README via
HuggingFaceService on selection change, and expose the download
progress type globally for the status feed. The download options
already read their state from the models status store.

Assisted-by: pi
Add a Discover models item to the selector dropdown footer that opens
the DialogModelsDiscover dialog.

Assisted-by: pi
- ModelsDiscoverItem + ModelsDiscoverInfo fold into
  ModelsDiscoverListItem (avatar + model id + badges + context/size)
- ModelsDiscoverDetails* renamed to ModelsDiscoverModelDetails*
- TerminalCommands renamed to ModelsDiscoverModelDetailsCommands
- ModelsDiscoverDetailsName folded into the details header
- ModelsDiscoverListSearch extracted from the list search input
- stories updated for the new names

Assisted-by: pi
- ModelsDiscoverListSearch: extracted search input from the list
- ModelsDiscoverModelDetailsMetadata: description + metadata chips
  extracted from the details header
- ModelsDiscoverModelDetailsCommands: quant + draft sidecar selectors
  embedded in the inline command text
- ModelsDownloadManager: tracked downloads with per-file progress and
  a delete action
- ModelsDownloadManagerDownloadStatusToast: one toast per download
  with a progress bar per file (main + sidecars) and a CTA to open
  the download manager
- DialogModelsDownloadManager: dialog shell for the manager

Assisted-by: pi
The quant and draft sidecar buttons become a multiple toggle group
(selecting which files to download, not download triggers), and below
the panel the terminal llama serve command updates live to reflect the
selection, with a download CTA that fires the downloads for the
selected entries (main + optional draft sidecar).

Adds the shadcn-svelte toggle-group component.

Assisted-by: pi
Downloaded files are no longer selectable toggles; they render as a
chip with a checkmark instead.

Assisted-by: pi
extractQuantMeta parsed the full sibling path, so repo layouts that nest
sidecars in a folder (e.g. MTP/mtp-Model-Q4_0.gguf) failed the sidecar
prefix match and were classified as main weights.

Assisted-by: pi:GLM-5.3-Flash
Replace the submenu-based dropdown with a flat scrollable list that has

a sticky search header and actions footer. Add avatars and parameter

badges to model options. Introduce ModelsSelectorReasoningPanel for

in-place reasoning effort selection. Adjust dropdown sizing and input

blur styling.

Assisted-by: llama-ui:Qwen3.8-Flash-Next
Move the expand chevron to the right and swap it for ChevronUp when

open. Group the effort label next to the title.

Assisted-by: llama-ui:Qwen3.8-Flash-Next
The inline picks now mirror the selection one-way instead of seeding $state, the default quant seeds once when the file list resolves, and the command only renders while something is selected.

Assisted-by: pi:zai-org/GLM-5.3
Assisted-by: pi:zai-org/GLM-5.3
Solo sidecar downloads register in /v1/models under the entry tag, so the state checks that first (normalized against the UD- quant prefix) and keeps the --model-draft / --mmproj args of registered models as a fallback. The chips no longer flip to downloaded while a download is in progress.

Assisted-by: pi:zai-org/GLM-5.3
…cars

A Q4_0-mtp style tag now resolves the sidecar file when no model file matches it, so a solo draft or mmproj download actually pulls the file. Cached sidecar files list as their own entries so the state survives a restart, and removing such a tag deletes only the sidecar.

Assisted-by: pi:zai-org/GLM-5.3
Entries like repo:Q4_0-mtp mark a downloaded sidecar file, not a loadable model.

Assisted-by: pi:zai-org/GLM-5.3
Replace the selector dropdown footer entry with a sidebar action that opens the discover dialog.

Assisted-by: pi:zai-org/GLM-5.3
Move the option's fallback logic into useModelParamsFallback and reuse it for the dropdown and sheet trigger badges.

Assisted-by: pi:zai-org/GLM-5.3
Hide the ModelId badge and icon wrappers when empty, normalize its indentation, and tighten the reasoning panel and discover detail spacing.

Assisted-by: pi:zai-org/GLM-5.3
Every quant chip is now an independent download action with its own
lifecycle: download, retry, pause, resume, cancel and delete with
confirmation. The store distinguishes user-stopped downloads over the
/models/sse feed so they settle silently, keeps paused progress
resumable, resolves the server-registered id on delete and refetches
the model list so the selector stays in sync with downloads.

Drops the selection toggle group, the download CTA and the download
manager dialogs and progress toasts. The serve command preview is now
a standalone widget with its own picks and an addable draft segment.

Assisted-by: pi
A "Download in progress" section above the loaded models lists the
in-flight and paused downloads with a live progress bar and the same
pause, resume and cancel actions as the discover quant chips, with the
same avatar and model id presentation as the regular option rows.

Assisted-by: pi
The download monitor thread acquires the mutex on its way out, so joining
it while holding the lock in server_models::remove deadlocks once the
status has flipped to DOWNLOADED. Join outside the lock, same pattern as
load_models().

Assisted-by: pi:zai-org/GLM-5.3
extractQuantMeta now recognizes a -draft tail after the sidecar token
(Model-MTP-draft.gguf), a bare sidecar token filename (imatrix.gguf) and
suffix sidecars whose head carries no quant (Model-imatrix.gguf). imatrix
is an aux sidecar: it stays in the download chips with a badge, is never
a serve-command option, and mmproj alone drives the Vision capability.

Assisted-by: pi:zai-org/GLM-5.3
Cancelling stops the download and discards the partial files, so the
selector item and the download chip both ask for confirmation first.

Assisted-by: pi:zai-org/GLM-5.3
Favorites sit above the in-flight downloads; the dropdown visual order
follows the new list order.

Assisted-by: pi:zai-org/GLM-5.3
Downloading entries are not usable models yet; the selector tracks them
in its download-progress section instead.

Assisted-by: pi:zai-org/GLM-5.3
The copy button swaps to a checkmark briefly; the close X moves into
the header row so it lines up with the title.

Assisted-by: pi:zai-org/GLM-5.3
Move the ModelsDiscover components into ModelsDiscoverList/ and
ModelsDiscoverDetails/ subfolders (with ModelsDiscoverDetailsDownloadOptions/
nested inside), and the model selector components into a ModelsSelector/
subfolder. Rename ModelsDiscoverModelDetails* to ModelsDiscoverDetails* for
consistency, add per-folder index.ts barrels, and keep shared leaves
(ModelId, ModelBadge, ModelLoadHighlight, ModelsDiscoverAvatar,
DownloadProgressBar) at their folder roots.

@allozaur allozaur left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review notes — memory-fit UI info

Goal for this area: keep only the same memory-fit UI logic as the llama.app models pages, and introduce no device-memory checking in our changes.

Good news — no device-memory logic remains. The device-tier machinery (resolveDeviceMemoryGb, deviceMemoryBudgetMb, computeFileCompatibilityTiers, CompatibilityTier, navigator.deviceMemory, the full/limited/none verdict) was dropped in #27957 (a07359ef) and never came back on the stack base.

One thing to reconcile — tools/ui/src/lib/utils/model-compatibility.ts currently exposes two functions:

  • estimateModelMemoryBytes() — the simplified size × 1.05; appears unused in the discover UI now.
  • minMemoryTierGb() — re-added on this branch (46c1d39b1); functionally identical to llama.app's ModelsCatalogService.minMemForBuild (same MAC_MEM_TIERS, RAM_BUDGET_RATIO = 0.75, RAM_OVERHEAD_MB = 2048, QUANT_WEIGHT = 1.05). It is a pure file-size → required-tier function (no device detection), rendered as "needs at least NGB+ memory" in the download-options row.

minMemoryTierGb is the keeper — it matches the llama.app behavior. Suggest dropping estimateModelMemoryBytes if it's truly unused, so we expose a single memory-fit function that mirrors minMemForBuild.

No action needed on device logic — it's already clean.

Comment thread tools/ui/src/lib/components/app/models/ModelId.svelte
Comment thread common/download.cpp
Comment thread tools/ui/tests/stories/ModelsSelector.stories.svelte
@allozaur
allozaur force-pushed the allozaur/ui-reuse-markers branch 2 times, most recently from d448211 to 1884398 Compare September 4, 2026 19:43
@allozaur

allozaur commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by the recomposed stack #28412: #28405 -> #28406 -> #27945 -> #27946 -> #27947 -> #27957 -> #27959 -> #28407 -> #28408 -> #28409 -> #28410 -> #28411 -> #28084. All review threads here are resolved and addressed in the corresponding layers.

@allozaur allozaur closed this Sep 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant