Skip to content

[Bug]: T3 writes origin.surface="cli" that its own event-store decoder rejects, making t3code unbootable #8789

Description

@Aryan-Saini

What happened

t3code.service became unbootable and crash-looped into start-limit-hit. The cause is a row in orchestration_events that T3 itself wrote, carrying an origin.surface value its own decoder rejects:

PersistenceDecodeError: Decode error in OrchestrationEventStore.readFromSequence:decodeRows
  Composite(Pointer(Composite(Pointer(Composite(Pointer(AnyOf(Composite(Pointer(AnyOf(AnyOf()))))))))))
  [cause]: SchemaError: Expected "web" | "desktop" | "mobile"
    at [0]["metadata"]["origin"]["surface"]

The offending row:

sequence:       43854
event_type:     project.meta-updated
actor_kind:     client
metadata_json:  {"origin":{"surface":"cli"}}

"cli" is not in the web | desktop | mobile enum the reader accepts, so readFromSequence throws during startup and the process exits 1. systemd retries, hits StartLimitBurst, and the service is down until someone intervenes.

Distribution across ~43.8k events on this machine — one bad row is enough:

origin.surface rows
absent 43,588
desktop 263
cli 1

Note that absent is overwhelmingly the norm and decodes fine, so the field is effectively optional; it is only the out-of-enum value that is fatal.

Why it is nasty

The event store is decoded only at startup. A poisoned row is completely harmless while the process is running, so:

  • The service stays green indefinitely — it had been up for days.
  • Health checks pass. Nothing surfaces a problem.
  • The failure lands on the next restart, from any cause: an auto-update, a reboot, a crash recovery.

So the blast radius is "the box does not come back", detonating at an arbitrary later time with no connection to whatever wrote the row. For a headless machine reached through T3 Connect, that means losing remote access entirely — the relay stays up and answers, but there is no backend behind it, surfacing to clients as Relay could not reach the environment endpoint (endpoint_request_failed).

It recurred: I cleared the first bad row, the service ran fine for about an hour, then a second project.meta-updated / origin.surface="cli" row appeared and the next restart crash-looped again.

Reproduction

I could not identify the writer, which is the main thing I would want help with.

Ruled out by direct test: t3 CLI reads. Running t3 connect status did not produce a bad row (count stayed 0 across the invocation).

What the two poisoned rows had in common:

  • event_type = project.meta-updated
  • actor_kind = client
  • occurred_at with exact .000Z millisecond precision, unlike surrounding rows, which suggests a different client/serialiser than the local service.

To check whether an install is affected:

SELECT json_extract(metadata_json,'$.origin.surface') AS surface, COUNT(*)
FROM orchestration_events
GROUP BY surface ORDER BY 2 DESC;

Suggested fix

The asymmetry is the bug: something in the write path emits "cli" while the read schema does not accept it. Either end would fix this instance, but the third point is what stops the whole class:

  1. Add cli to the origin.surface enum, if CLI-origin events are legitimate.
  2. Otherwise stop the writer emitting it.
  3. Make readFromSequence resilient to undecodable rows. A single malformed event in an append-only log should not be able to render the service unbootable. Skipping, quarantining or lenient-decoding a bad row with a loud warning would turn a hard brick into a logged anomaly.

This is the same shape as #4518 (persisted value rejected by the startup decoder, backend crash-loops). That was a settings value; this one is an event row, so it can be introduced by normal operation rather than by a user changing a setting.

Workaround

For anyone hitting this — clear the out-of-enum value and reset the latched start limit:

UPDATE orchestration_events
SET metadata_json = json_remove(metadata_json, '$.origin.surface')
WHERE json_extract(metadata_json, '$.origin.surface') IS NOT NULL
  AND json_extract(metadata_json, '$.origin.surface') NOT IN ('web','desktop','mobile');
systemctl --user reset-failed t3code.service   # the start limit latches and blocks a plain restart
systemctl --user start t3code.service

Because it recurs, I ended up guarding the start path so the box can self-recover:

ExecStartPre=-%h/.local/bin/t3-sanitize-event-store

Environment

  • t3 0.0.37-nightly.20260829.1219 (nightly channel)
  • Linux (Arch), headless server, node 24.19.0
  • Run as a systemd --user service with Restart=always, StartLimitBurst=5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions