Skip to content

boot: schema discovery should not be fatal — retry with backoff, serve diagnostic /health 503 #95

Description

@jfwoods

Problem

In cmd/wavehouse standalone mode, the schema-discovery path (main.go:114) calls
os.Exit on any failure. Two real triggers seen in deployment:

  • Missing database: WH_CH_DATABASE pointing at a database never created on
    that ClickHouse instance → schema discovery failed on boot — query system.columns: code: 81, message: Database <db> does not exist. Nothing in the
    stack applies wavehouse/schemas/*.sql automatically, so the operator must do it
    by hand.
  • Transient unreachability: ClickHouse briefly not accepting connections on 9000
    (e.g. a compose change shadowing CH's stock docker_related_config.xml, which
    sets <listen_host>::</listen_host>) → dial tcp …:9000: connect: connection refused.

In both cases WaveHouse exits, the supervisor restarts it ~every 10s in an
unbounded loop, port 8080 never binds, and clients get connection refused even
though DNS resolves and (case 2) ClickHouse is otherwise healthy. The binary is
unrecoverable without operator intervention.

Proposed Solution

Schema discovery on boot should never be fatal. Connection-refused,
missing-database, and any other transient/configuration error should trigger a
backoff-retry instead of process exit. While discovery is failing, the process
should still bind :8080 and serve /health 503 with the diagnostic message, so
an operator can curl /health instead of grepping a restart-loop log. Once
ClickHouse is reachable and the schema resolves, /health flips to 200 and normal
serving begins.

Alternatives Considered

  • Auto-apply the schema DDL on boot when the database is missing — too
    magical/destructive for a gateway; out of scope here.
  • Add a longer fixed retry/timeout before exit — still leaves the binary in a
    restart loop; doesn't give operators a queryable health surface.

Additional Context

From nas-observability WHissues.md, 2026-05-06 and 2026-05-08. Adjacent to #46
(Graceful Shutdown), but that covers shutdown, not boot; and parallel to #49 (NATS
Robustness & Backoff) — same "don't crash-loop, back off, log degraded mode"
philosophy applied to the ClickHouse-on-boot path instead of NATS reconnect.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area/infraCI, build, deploy, Docker, releasearea/ingestIngest pipeline (Bento, batching, DLQ)area/observabilityMetrics, logs, traces, health, profilingbugSomething isn't workingenhancementNew feature or request

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions