Problem
In cmd/wavehouse standalone mode, the schema-discovery path (main.go:114) calls
os.Exit on any failure. Two real triggers seen in deployment:
- Missing database:
WH_CH_DATABASE pointing at a database never created on
that ClickHouse instance → schema discovery failed on boot — query system.columns: code: 81, message: Database <db> does not exist. Nothing in the
stack applies wavehouse/schemas/*.sql automatically, so the operator must do it
by hand.
- Transient unreachability: ClickHouse briefly not accepting connections on 9000
(e.g. a compose change shadowing CH's stock docker_related_config.xml, which
sets <listen_host>::</listen_host>) → dial tcp …:9000: connect: connection refused.
In both cases WaveHouse exits, the supervisor restarts it ~every 10s in an
unbounded loop, port 8080 never binds, and clients get connection refused even
though DNS resolves and (case 2) ClickHouse is otherwise healthy. The binary is
unrecoverable without operator intervention.
Proposed Solution
Schema discovery on boot should never be fatal. Connection-refused,
missing-database, and any other transient/configuration error should trigger a
backoff-retry instead of process exit. While discovery is failing, the process
should still bind :8080 and serve /health 503 with the diagnostic message, so
an operator can curl /health instead of grepping a restart-loop log. Once
ClickHouse is reachable and the schema resolves, /health flips to 200 and normal
serving begins.
Alternatives Considered
- Auto-apply the schema DDL on boot when the database is missing — too
magical/destructive for a gateway; out of scope here.
- Add a longer fixed retry/timeout before exit — still leaves the binary in a
restart loop; doesn't give operators a queryable health surface.
Additional Context
From nas-observability WHissues.md, 2026-05-06 and 2026-05-08. Adjacent to #46
(Graceful Shutdown), but that covers shutdown, not boot; and parallel to #49 (NATS
Robustness & Backoff) — same "don't crash-loop, back off, log degraded mode"
philosophy applied to the ClickHouse-on-boot path instead of NATS reconnect.
Problem
In
cmd/wavehousestandalone mode, the schema-discovery path (main.go:114) callsos.Exiton any failure. Two real triggers seen in deployment:WH_CH_DATABASEpointing at a database never created onthat ClickHouse instance →
schema discovery failed on boot — query system.columns: code: 81, message: Database <db> does not exist. Nothing in thestack applies
wavehouse/schemas/*.sqlautomatically, so the operator must do itby hand.
(e.g. a compose change shadowing CH's stock
docker_related_config.xml, whichsets
<listen_host>::</listen_host>) →dial tcp …:9000: connect: connection refused.In both cases WaveHouse exits, the supervisor restarts it ~every 10s in an
unbounded loop, port 8080 never binds, and clients get
connection refusedeventhough DNS resolves and (case 2) ClickHouse is otherwise healthy. The binary is
unrecoverable without operator intervention.
Proposed Solution
Schema discovery on boot should never be fatal. Connection-refused,
missing-database, and any other transient/configuration error should trigger a
backoff-retry instead of process exit. While discovery is failing, the process
should still bind
:8080and serve/health503 with the diagnostic message, soan operator can
curl /healthinstead of grepping a restart-loop log. OnceClickHouse is reachable and the schema resolves,
/healthflips to 200 and normalserving begins.
Alternatives Considered
magical/destructive for a gateway; out of scope here.
restart loop; doesn't give operators a queryable health surface.
Additional Context
From nas-observability
WHissues.md, 2026-05-06 and 2026-05-08. Adjacent to #46(Graceful Shutdown), but that covers shutdown, not boot; and parallel to #49 (NATS
Robustness & Backoff) — same "don't crash-loop, back off, log degraded mode"
philosophy applied to the ClickHouse-on-boot path instead of NATS reconnect.