Repository navigation
Conversation
ApprovabilityVerdict: Approved at Macroscope's review found this PR approvable — This is a focused connection-supervisor bug fix: desktop/web foreground probes retry once before reconnecting, while mobile and explicit-failure paths remain unchanged. The behavior is covered by targeted tests, and the change introduces no schema, infrastructure, security, billing, or static-analysis configuration impact. Notes:
You can add or adjust custom eligibility rules. Learn more. |
| Effect.annotateLogs({ | ||
| "environment.id": target.environmentId, | ||
| "environment.label": target.label, | ||
| }), |
There was a problem hiding this comment.
The new warning logs target.label verbatim, but connection labels are unrestricted strings. Please keep log annotations bounded and safe by recording the environment ID without the label.
| Effect.annotateLogs({ | |
| "environment.id": target.environmentId, | |
| "environment.label": target.label, | |
| }), | |
| Effect.annotateLogs({ | |
| "environment.id": target.environmentId, | |
| }), |
Posted via Macroscope — Effect Service Conventions
|
Correction to my inline suggestion on Posted via Macroscope — Effect Service Conventions |
1164990 to
098103f
Compare
|
Note Posted by an AI agent on behalf of Tyler. Rebased onto current main. The timeout warning now only annotates |
Fixes #7231.
Problem
When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".
The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.
Fix
On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.
The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:
Fiber.interrupt(probe)paths.application-activewakeup to decide. On desktop that wakeup only fires onvisibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.Switching the desktop path to
Effect.timeoutOptionalso removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives asNoneon the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.Mobile is unchanged.
application-active-probekeeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bareapplication-activereason, so the retry path is unreachable from it.Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through
session.closed, so nothing here is the sole defence against a dropped socket.Relationship to #3553, #4137, and #5198
#3553 reported this and was closed by #4137, which made the probe itself lightweight (
serverProberather thanserverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.
Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.
Testing
typecheckandvp lint packages/client-runtime/src/connectionclean.Docs:
docs/internals/connection-runtime.md"Wakeups" section now states the retry, the bound, and the mobile exclusion.Model: Claude Opus 5; harness: Claude Code.
Note
Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.
Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile
application-active-probestays 3s fail-fast with no retry.monitorConnectedLeaseswitches the desktop path toEffect.timeoutOptionand runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existingFiber.interruptpaths, with worst-case recovery bounded (~35s).Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.
Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.
Note
Fix desktop/web foreground probe to retry once before treating a stall as a disconnect
CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.Macroscope summarized 1164990.