Skip to content

Misc. bug: llama-server --models-preset creates an unexpected "default" model entry #22364

Description

@vanbukin

Name and Version

root@pc2:~# /opt/llama.cpp/bin/llama-server --version
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 24079 MiB):
  Device 0: NVIDIA GeForce RTX 4090, compute capability 8.9, VMM: yes, VRAM: 24079 MiB
version: 8931 (9725a313b)
built with GNU 15.2.0 for Linux x86_64
root@pc2:~# cat /etc/os-release
PRETTY_NAME="Ubuntu 26.04 LTS"
NAME="Ubuntu"
VERSION_ID="26.04"
VERSION="26.04 (Resolute Raccoon)"
VERSION_CODENAME=resolute
ID=ubuntu
ID_LIKE=debian
HOME_URL="https://www.ubuntu.com/"
SUPPORT_URL="https://help.ubuntu.com/"
BUG_REPORT_URL="https://bugs.launchpad.net/ubuntu/"
PRIVACY_POLICY_URL="https://www.ubuntu.com/legal/terms-and-policies/privacy-policy"
UBUNTU_CODENAME=resolute
LOGO=ubuntu-logo

Operating systems

Linux

Which llama.cpp modules do you know to be affected?

llama-server

Command line

/opt/llama.cpp/bin/llama-server \
--host 127.0.0.1 \
--port 8080 \
--models-preset /etc/llama.cpp/models.ini \
--models-max 1 \
--models-autoload \
--api-key-file /etc/llama.cpp/api_keys \
--metrics

Problem description & steps to reproduce

When running llama-server in router mode with --models-preset, an extra model with id default appears in GET /models, even though there is no [default] section in the preset file and no model is specified on the command line.

According to the documentation, each section in the INI file defines a new preset. In my case, the INI file contains only two model sections, but the server reports three custom presets and exposes a third default model.

Preset file

/etc/llama.cpp/models.ini:

version = 1

[qwen3.6-27b]
model = /srv/llama.cpp/models/unsloth/Qwen3.6-27B-GGUF/Qwen3.6-27B-UD-Q4_K_XL.gguf
mmproj = /srv/llama.cpp/models/unsloth/Qwen3.6-27B-GGUF/mmproj-BF16.gguf
ctx-size = 122880
ubatch-size = 1024
batch-size = 122880
n-cpu-moe = 0
n-gpu-layers = 999
temp = 0.6
top-p = 0.95
top-k = 20
min-p = 0.00
repeat-penalty = 1.0
presence-penalty = 0.0
split-mode = none
main-gpu = 0
image-min-tokens = 1024
image-max-tokens = 2048
parallel = 1
flash-attn = on
cache-type-k = q8_0
cache-type-v = q8_0
chat-template-kwargs = {"enable_thinking":true}
load-on-startup = true
stop-timeout = 10

[qwen3-asr-1.7b]
model = /srv/llama.cpp/models/ggml-org/Qwen3-ASR-1.7B-GGUF/Qwen3-ASR-1.7B-bf16.gguf
mmproj = /srv/llama.cpp/models/ggml-org/Qwen3-ASR-1.7B-GGUF/mmproj-Qwen3-ASR-1.7B-bf16.gguf
ctx-size = 65536
ubatch-size = 2048
batch-size = 65536
n-cpu-moe = 0
n-gpu-layers = 999
split-mode = none
main-gpu = 0
parallel = 1
load-on-startup = false
stop-timeout = 10

Steps to reproduce

  1. Create a model preset file with two model sections and no [default] section, as shown above.
  2. Start llama-server in router mode:
/opt/llama.cpp/bin/llama-server \
--host 127.0.0.1 \
--port 8080 \
--models-preset /etc/llama.cpp/models.ini \
--models-max 1 \
--models-autoload \
--api-key-file /etc/llama.cpp/api_keys \
--metrics
  1. Query the models endpoint:
curl http://127.0.0.1:8080/models

Expected behavior

Only the two models explicitly defined in the preset file should be listed:

  • qwen3-asr-1.7b
  • qwen3.6-27b

The startup log should also report 2 custom model presets loaded from /etc/llama.cpp/models.ini.

Actual behavior

The server reports 3 custom model presets and exposes an unexpected third model named default

{
    "data": [
        {
            "id": "default",
            "aliases": [],
            "tags": [],
            "object": "model",
            "owned_by": "llamacpp",
            "created": 1777137717,
            "status": {
                "value": "unloaded",
                "args": [
                    "/opt/llama.cpp/bin/llama-server",
                    "--api-key-file",
                    "/etc/llama.cpp/api_keys",
                    "--host",
                    "127.0.0.1",
                    "--metrics",
                    "--port",
                    "0",
                    "--alias",
                    "default"
                ],
                "preset": "[default]\napi-key-file = /etc/llama.cpp/api_keys\nmetrics = 1\n\n"
            }
        },
        {
            "id": "qwen3-asr-1.7b",
            "aliases": [],
            "tags": [],
            "object": "model",
            "owned_by": "llamacpp",
            "created": 1777137717,
            "status": {
                "value": "unloaded",
                "args": [
                    "/opt/llama.cpp/bin/llama-server",
                    "--api-key-file",
                    "/etc/llama.cpp/api_keys",
                    "--host",
                    "127.0.0.1",
                    "--metrics",
                    "--port",
                    "0",
                    "--alias",
                    "qwen3-asr-1.7b",
                    "--batch-size",
                    "65536",
                    "--ctx-size",
                    "65536",
                    "--model",
                    "/srv/llama.cpp/models/ggml-org/Qwen3-ASR-1.7B-GGUF/Qwen3-ASR-1.7B-bf16.gguf",
                    "--main-gpu",
                    "0",
                    "--mmproj",
                    "/srv/llama.cpp/models/ggml-org/Qwen3-ASR-1.7B-GGUF/mmproj-Qwen3-ASR-1.7B-bf16.gguf",
                    "--n-cpu-moe",
                    "0",
                    "--n-gpu-layers",
                    "999",
                    "--parallel",
                    "1",
                    "--split-mode",
                    "none",
                    "--ubatch-size",
                    "2048"
                ],
                "preset": "[qwen3-asr-1.7b]\napi-key-file = /etc/llama.cpp/api_keys\nmetrics = 1\nbatch-size = 65536\nctx-size = 65536\nmodel = /srv/llama.cpp/models/ggml-org/Qwen3-ASR-1.7B-GGUF/Qwen3-ASR-1.7B-bf16.gguf\nmain-gpu = 0\nmmproj = /srv/llama.cpp/models/ggml-org/Qwen3-ASR-1.7B-GGUF/mmproj-Qwen3-ASR-1.7B-bf16.gguf\nn-cpu-moe = 0\nn-gpu-layers = 999\nparallel = 1\nsplit-mode = none\nubatch-size = 2048\nload-on-startup = false\nstop-timeout = 10\n\n"
            }
        },
        {
            "id": "qwen3.6-27b",
            "aliases": [],
            "tags": [],
            "object": "model",
            "owned_by": "llamacpp",
            "created": 1777137717,
            "status": {
                "value": "loaded",
                "args": [
                    "/opt/llama.cpp/bin/llama-server",
                    "--api-key-file",
                    "/etc/llama.cpp/api_keys",
                    "--chat-template-kwargs",
                    "{\"enable_thinking\":true}",
                    "--host",
                    "127.0.0.1",
                    "--image-max-tokens",
                    "2048",
                    "--image-min-tokens",
                    "1024",
                    "--metrics",
                    "--min-p",
                    "0.00",
                    "--port",
                    "53891",
                    "--presence-penalty",
                    "0.0",
                    "--repeat-penalty",
                    "1.0",
                    "--temperature",
                    "0.6",
                    "--top-k",
                    "20",
                    "--top-p",
                    "0.95",
                    "--alias",
                    "qwen3.6-27b",
                    "--batch-size",
                    "122880",
                    "--ctx-size",
                    "122880",
                    "--cache-type-k",
                    "q8_0",
                    "--cache-type-v",
                    "q8_0",
                    "--flash-attn",
                    "on",
                    "--model",
                    "/srv/llama.cpp/models/unsloth/Qwen3.6-27B-GGUF/Qwen3.6-27B-UD-Q4_K_XL.gguf",
                    "--main-gpu",
                    "0",
                    "--mmproj",
                    "/srv/llama.cpp/models/unsloth/Qwen3.6-27B-GGUF/mmproj-BF16.gguf",
                    "--n-cpu-moe",
                    "0",
                    "--n-gpu-layers",
                    "999",
                    "--parallel",
                    "1",
                    "--split-mode",
                    "none",
                    "--ubatch-size",
                    "1024"
                ],
                "preset": "[qwen3.6-27b]\napi-key-file = /etc/llama.cpp/api_keys\nchat-template-kwargs = {\"enable_thinking\":true}\nimage-max-tokens = 2048\nimage-min-tokens = 1024\nmetrics = 1\nmin-p = 0.00\npresence-penalty = 0.0\nrepeat-penalty = 1.0\ntemperature = 0.6\ntop-k = 20\ntop-p = 0.95\nbatch-size = 122880\nctx-size = 122880\ncache-type-k = q8_0\ncache-type-v = q8_0\nflash-attn = on\nmodel = /srv/llama.cpp/models/unsloth/Qwen3.6-27B-GGUF/Qwen3.6-27B-UD-Q4_K_XL.gguf\nmain-gpu = 0\nmmproj = /srv/llama.cpp/models/unsloth/Qwen3.6-27B-GGUF/mmproj-BF16.gguf\nn-cpu-moe = 0\nn-gpu-layers = 999\nparallel = 1\nsplit-mode = none\nubatch-size = 1024\nload-on-startup = true\nstop-timeout = 10\n\n"
            }
        }
    ],
    "object": "list"
}

The unexpected default entry is not defined in the INI file and does not have a model path.

This is confusing for clients that use GET /models to discover available models, because the endpoint returns a model that was never configured and cannot be used as a real preset.

First Bad Commit

No response

Relevant log output

Startup log

Logs
root@pc2:~# /opt/llama.cpp/bin/llama-server --host 127.0.0.1 --port 8080 --models-preset /etc/llama.cpp/models.ini --models-max 1 --models-autoload --api-key-file /etc/llama.cpp/api_keys --metrics
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 24079 MiB):
  Device 0: NVIDIA GeForce RTX 4090, compute capability 8.9, VMM: yes, VRAM: 24079 MiB
main: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
build_info: b8931-9725a313b
system_info: n_threads = 16 (n_threads_batch = 16) / 32 | CUDA : ARCHS = 890 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
Running without SSL
init: api_keys: ****42a8
init: using 31 threads for HTTP server
srv   load_models: Loaded 0 cached model presets
srv   load_models: Loaded 3 custom model presets from /etc/llama.cpp/models.ini
srv   load_models: Available models (3) (*: custom preset)
srv   load_models:   * default
srv   load_models:   * qwen3-asr-1.7b
srv   load_models:   * qwen3.6-27b
srv   load_models: (startup) loading model qwen3.6-27b
srv          load: spawning server instance with name=qwen3.6-27b on port 53891
...

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions