What I see generally:
- An unreachable virtual model provider causes a crash, no virtual models, it starts up successfully with the same down provider
- The current implementation has an incorrect/inconsistent error, making
ModelRegistry.InitializeAsync synchronous shows the real error
- The expected behavior I would like to see is (a) don't crash, status reflected (b) a log message when this init/validation process completes. Maybe this should not even be validating that the providers via virtual models are already responding? I had an agent do something similar in an
Init style function. It merged the initial "config->object in memory" process with checking if the llm providers were available, distinct processes in my opinion because config loading and object creation should not involve API calls.
Config
(relevant)
virtual_models:
- source: qwen-coder
targets:
- provider: uranium-8000
model: lovedheart/Qwen-AgentWorld-35B-A3B-NVFP4
- provider: plutonium-8000
model: lovedheart/Qwen-AgentWorld-35B-A3B-NVFP4
- source: gemma-general
targets:
- provider: plutonium-8001
model: nvidia/Gemma-4-26B-A4B-NVFP4
- source: embedding
targets:
- provider: uranium-8001
model: google/embeddinggemma-300m
- source: microllm
targets:
- provider: uranium-8002
model: Qwen/Qwen3-1.7B
- source: reranker
targets:
- provider: uranium-8003
model: Qwen/Qwen3-Reranker-0.6B
models:
enabled_by_default: true # env: MODELS_ENABLED_BY_DEFAULT; when false, models stay unavailable until an access override allows one or more user paths
configured_provider_models_mode: "fallback" # env: CONFIGURED_PROVIDER_MODELS_MODE; "fallback" uses configured lists only when upstream /models is unavailable/empty, "allowlist" exposes only configured models and skips upstream /models for configured lists
providers:
uranium-8000:
type: vllm
base_url: "http://192.168.4.31:8000/v1"
uranium-8001:
type: vllm
base_url: "http://192.168.4.31:8001/v1"
uranium-8002:
type: vllm
base_url: "http://192.168.4.31:8002/v1"
uranium-8003:
type: vllm
base_url: "http://192.168.4.31:8003/v1"
plutonium-8000:
type: vllm
base_url: "http://192.168.4.35:8000/v1"
plutonium-8001:
type: vllm
base_url: "http://192.168.4.35:8001/v1"
# Global resilience settings (applied to all providers by default)
# Individual providers can override any of these values.
resilience:
retry:
max_retries: 3
initial_backoff: 6s
max_backoff: 30s
backoff_factor: 2.0
jitter_factor: 0.1
With Async Registry Init (current)
This is both fast and inaccurate. It seems like if any provider is not reachable, then the InitializeAsync (registry_init.go) doesn't finish processing the initial population, and it randomly reports which model was the issue.
+ docker run --rm -it --name gomodel -p 9999:9999 -v /Users/tony/verdverm/gmd/models/gomodel/config.yaml:/app/config/config.yaml:ro enterpilot/gomodel:local
11:57PM INF starting gomodel version=dev commit=none build_date=unknown
11:57PM INF using local file cache path=.cache/models.json
11:57PM INF provider registered name=plutonium-8000 type=vllm
11:57PM INF provider registered name=plutonium-8001 type=vllm
11:57PM INF provider registered name=uranium-8000 type=vllm
11:57PM INF provider registered name=uranium-8001 type=vllm
11:57PM INF provider registered name=uranium-8002 type=vllm
11:57PM INF provider registered name=uranium-8003 type=vllm
11:57PM INF starting non-blocking model registry initialization...
11:57PM INF model registry configured cached_models=0 providers=6
11:57PM INF model list loaded models=778 providers=19 provider_models=1836 metadata_enriched=0 metadata_total=0 metadata_providers=0
11:57PM ERR failed to initialize application error="failed to initialize virtual models: load virtual model \"embedding\": target model not found: uranium-8001/google/embeddinggemma-300m"
With Sync Registry Init (modified)
I removed the go func() { ... }() and it fails with a better message and reports the correct model it failed to load.
+ docker run --rm -it --name gomodel -p 9999:9999 -v /Users/tony/verdverm/gmd/models/gomodel/config.yaml:/app/config/config.yaml:ro enterpilot/gomodel:local
11:52PM INF starting gomodel version=dev commit=none build_date=unknown
11:52PM INF using local file cache path=.cache/models.json
11:52PM INF provider registered name=plutonium-8000 type=vllm
11:52PM INF provider registered name=plutonium-8001 type=vllm
11:52PM INF provider registered name=uranium-8000 type=vllm
11:52PM INF provider registered name=uranium-8001 type=vllm
11:52PM INF provider registered name=uranium-8002 type=vllm
11:52PM INF provider registered name=uranium-8003 type=vllm
11:52PM INF starting non-blocking model registry initialization...
11:53PM WRN failed to fetch models from provider provider=plutonium-8001 error="[vllm] provider_error: failed to send request: Get \"http://192.168.4.35:8001/v1/models\": dial tcp 192.168.4.35:8001: connect: connection refused"
11:53PM INF model registry initialized total_models=4 providers=6 failed_providers=1 metadata_enriched=0 metadata_total=0 metadata_providers=0
11:53PM INF model registry configured cached_models=4 providers=6
11:53PM ERR failed to initialize application error="failed to initialize virtual models: load virtual model \"gemma-general\": target model not found: plutonium-8001/nvidia/Gemma-4-26B-A4B-NVFP4"
Without Virtual Models
This starts up and eventually shows the last line and failed provider.
+ docker run --rm -it --name gomodel -p 9999:9999 -v /Users/tony/verdverm/gmd/models/gomodel/config.yaml:/app/config/config.yaml:ro enterpilot/gomodel:local
12:08AM INF starting gomodel version=dev commit=none build_date=unknown
12:08AM INF using local file cache path=.cache/models.json
12:08AM INF provider registered name=plutonium-8000 type=vllm
12:08AM INF provider registered name=plutonium-8001 type=vllm
12:08AM INF provider registered name=uranium-8000 type=vllm
12:08AM INF provider registered name=uranium-8001 type=vllm
12:08AM INF provider registered name=uranium-8002 type=vllm
12:08AM INF provider registered name=uranium-8003 type=vllm
12:08AM INF starting non-blocking model registry initialization...
12:08AM INF model registry configured cached_models=0 providers=6
12:08AM WRN SECURITY WARNING: GOMODEL_MASTER_KEY not set - server running in UNSAFE MODE security_risk="unauthenticated access allowed" recommendation="set GOMODEL_MASTER_KEY environment variable to secure this gateway"
12:08AM INF prometheus metrics enabled endpoint=/metrics
12:08AM INF storage configured type=sqlite
12:08AM INF audit logging enabled log_bodies=true log_audio_bodies=false log_headers=true retention_days=30
12:08AM INF usage tracking enabled buffer_size=1000 flush_interval=5 retention_days=90
12:08AM INF admin API enabled api=/admin legacy_alias=/admin/api/v1 legacy_sunset=2026-08-09
12:08AM INF admin UI enabled url=http://localhost:9999/admin/dashboard
12:08AM INF provider passthrough enabled path=/p/{provider}/{endpoint}
12:08AM INF starting server address=:9999
12:08AM INF http(s) server started address=[::]:9999
12:08AM INF model list loaded models=778 providers=19 provider_models=1836 metadata_enriched=0 metadata_total=0 metadata_providers=0
12:09AM WRN failed to fetch models from provider provider=plutonium-8001 error="[vllm] provider_error: failed to send request: Get \"http://192.168.4.35:8001/v1/models\": dial tcp 192.168.4.35:8001: connect: connection refused"
12:09AM INF model registry initialized total_models=4 providers=6 failed_providers=1 metadata_enriched=0 metadata_total=5 metadata_providers=5
What I see generally:
ModelRegistry.InitializeAsyncsynchronous shows the real errorInitstyle function. It merged the initial "config->object in memory" process with checking if the llm providers were available, distinct processes in my opinion because config loading and object creation should not involve API calls.Config
(relevant)
With Async Registry Init (current)
This is both fast and inaccurate. It seems like if any provider is not reachable, then the InitializeAsync (registry_init.go) doesn't finish processing the initial population, and it randomly reports which model was the issue.
With Sync Registry Init (modified)
I removed the
go func() { ... }()and it fails with a better message and reports the correct model it failed to load.Without Virtual Models
This starts up and eventually shows the last line and failed provider.