What happened
t3code.service became unbootable and crash-looped into start-limit-hit. The cause is a row in orchestration_events that T3 itself wrote, carrying an origin.surface value its own decoder rejects:
PersistenceDecodeError: Decode error in OrchestrationEventStore.readFromSequence:decodeRows
Composite(Pointer(Composite(Pointer(Composite(Pointer(AnyOf(Composite(Pointer(AnyOf(AnyOf()))))))))))
[cause]: SchemaError: Expected "web" | "desktop" | "mobile"
at [0]["metadata"]["origin"]["surface"]
The offending row:
sequence: 43854
event_type: project.meta-updated
actor_kind: client
metadata_json: {"origin":{"surface":"cli"}}
"cli" is not in the web | desktop | mobile enum the reader accepts, so readFromSequence throws during startup and the process exits 1. systemd retries, hits StartLimitBurst, and the service is down until someone intervenes.
Distribution across ~43.8k events on this machine — one bad row is enough:
origin.surface |
rows |
| absent |
43,588 |
desktop |
263 |
cli |
1 |
Note that absent is overwhelmingly the norm and decodes fine, so the field is effectively optional; it is only the out-of-enum value that is fatal.
Why it is nasty
The event store is decoded only at startup. A poisoned row is completely harmless while the process is running, so:
- The service stays green indefinitely — it had been up for days.
- Health checks pass. Nothing surfaces a problem.
- The failure lands on the next restart, from any cause: an auto-update, a reboot, a crash recovery.
So the blast radius is "the box does not come back", detonating at an arbitrary later time with no connection to whatever wrote the row. For a headless machine reached through T3 Connect, that means losing remote access entirely — the relay stays up and answers, but there is no backend behind it, surfacing to clients as Relay could not reach the environment endpoint (endpoint_request_failed).
It recurred: I cleared the first bad row, the service ran fine for about an hour, then a second project.meta-updated / origin.surface="cli" row appeared and the next restart crash-looped again.
Reproduction
I could not identify the writer, which is the main thing I would want help with.
Ruled out by direct test: t3 CLI reads. Running t3 connect status did not produce a bad row (count stayed 0 across the invocation).
What the two poisoned rows had in common:
event_type = project.meta-updated
actor_kind = client
occurred_at with exact .000Z millisecond precision, unlike surrounding rows, which suggests a different client/serialiser than the local service.
To check whether an install is affected:
SELECT json_extract(metadata_json,'$.origin.surface') AS surface, COUNT(*)
FROM orchestration_events
GROUP BY surface ORDER BY 2 DESC;
Suggested fix
The asymmetry is the bug: something in the write path emits "cli" while the read schema does not accept it. Either end would fix this instance, but the third point is what stops the whole class:
- Add
cli to the origin.surface enum, if CLI-origin events are legitimate.
- Otherwise stop the writer emitting it.
- Make
readFromSequence resilient to undecodable rows. A single malformed event in an append-only log should not be able to render the service unbootable. Skipping, quarantining or lenient-decoding a bad row with a loud warning would turn a hard brick into a logged anomaly.
This is the same shape as #4518 (persisted value rejected by the startup decoder, backend crash-loops). That was a settings value; this one is an event row, so it can be introduced by normal operation rather than by a user changing a setting.
Workaround
For anyone hitting this — clear the out-of-enum value and reset the latched start limit:
UPDATE orchestration_events
SET metadata_json = json_remove(metadata_json, '$.origin.surface')
WHERE json_extract(metadata_json, '$.origin.surface') IS NOT NULL
AND json_extract(metadata_json, '$.origin.surface') NOT IN ('web','desktop','mobile');
systemctl --user reset-failed t3code.service # the start limit latches and blocks a plain restart
systemctl --user start t3code.service
Because it recurs, I ended up guarding the start path so the box can self-recover:
ExecStartPre=-%h/.local/bin/t3-sanitize-event-store
Environment
- t3
0.0.37-nightly.20260829.1219 (nightly channel)
- Linux (Arch), headless server, node 24.19.0
- Run as a
systemd --user service with Restart=always, StartLimitBurst=5
What happened
t3code.servicebecame unbootable and crash-looped intostart-limit-hit. The cause is a row inorchestration_eventsthat T3 itself wrote, carrying anorigin.surfacevalue its own decoder rejects:The offending row:
"cli"is not in theweb | desktop | mobileenum the reader accepts, soreadFromSequencethrows during startup and the process exits 1. systemd retries, hitsStartLimitBurst, and the service is down until someone intervenes.Distribution across ~43.8k events on this machine — one bad row is enough:
origin.surfacedesktopcliNote that
absentis overwhelmingly the norm and decodes fine, so the field is effectively optional; it is only the out-of-enum value that is fatal.Why it is nasty
The event store is decoded only at startup. A poisoned row is completely harmless while the process is running, so:
So the blast radius is "the box does not come back", detonating at an arbitrary later time with no connection to whatever wrote the row. For a headless machine reached through T3 Connect, that means losing remote access entirely — the relay stays up and answers, but there is no backend behind it, surfacing to clients as
Relay could not reach the environment endpoint (endpoint_request_failed).It recurred: I cleared the first bad row, the service ran fine for about an hour, then a second
project.meta-updated/origin.surface="cli"row appeared and the next restart crash-looped again.Reproduction
I could not identify the writer, which is the main thing I would want help with.
Ruled out by direct test: t3 CLI reads. Running
t3 connect statusdid not produce a bad row (count stayed 0 across the invocation).What the two poisoned rows had in common:
event_type = project.meta-updatedactor_kind = clientoccurred_atwith exact.000Zmillisecond precision, unlike surrounding rows, which suggests a different client/serialiser than the local service.To check whether an install is affected:
Suggested fix
The asymmetry is the bug: something in the write path emits
"cli"while the read schema does not accept it. Either end would fix this instance, but the third point is what stops the whole class:clito theorigin.surfaceenum, if CLI-origin events are legitimate.readFromSequenceresilient to undecodable rows. A single malformed event in an append-only log should not be able to render the service unbootable. Skipping, quarantining or lenient-decoding a bad row with a loud warning would turn a hard brick into a logged anomaly.This is the same shape as #4518 (persisted value rejected by the startup decoder, backend crash-loops). That was a settings value; this one is an event row, so it can be introduced by normal operation rather than by a user changing a setting.
Workaround
For anyone hitting this — clear the out-of-enum value and reset the latched start limit:
systemctl --user reset-failed t3code.service # the start limit latches and blocks a plain restart systemctl --user start t3code.serviceBecause it recurs, I ended up guarding the start path so the box can self-recover:
ExecStartPre=-%h/.local/bin/t3-sanitize-event-storeEnvironment
0.0.37-nightly.20260829.1219(nightly channel)systemd --userservice withRestart=always,StartLimitBurst=5