Skip to content

[noqa] fix: reconnect Pusher when a PONG is missed instead of only logging - #96883

Merged
mountiny merged 5 commits into
Expensify:mainfrom
callstack-internal:fix/pusher-pong-reconnect
Jul 28, 2026
Merged

mountiny merged 5 commits into
Expensify:mainfrom
callstack-internal:fix/pusher-pong-reconnect

Conversation

@adhorodyski

@adhorodyski adhorodyski commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

Explanation of Change

The app sends a PING to Pusher every 30 seconds and expects a PONG back. A watchdog checks for missing PONGs every 60 seconds. Today, when the PONG never comes back (for example after the machine sleeps, or on a network that throttles websockets), the watchdog only writes a misleading "going offline" log line and does nothing — so chat and report updates silently stop until the user reloads the app.

Now the watchdog calls Pusher.reconnect() when the PONG has been missing for over 60 seconds. Each reconnect resets the PONG clock, so the new socket gets a fresh grace period and retries space out naturally to about one reconnect every 2 minutes. Retries continue for as long as PONGs stay missing and stop the moment one arrives. Reconnecting also re-subscribes the user channel, which triggers reconnectApp, so the messages missed while the socket was dead are fetched too. The misleading log text is fixed.

Fixed Issues

$ #93966
PROPOSAL:

Tests

  1. Run npx jest tests/unit/PusherPingPongTest.ts — 2 tests pass: the watchdog reconnects once the PONG goes missing and keeps retrying about every 2 minutes while PONGs stay missing, and a fresh PONG defers the next reconnect.
  2. Manual: put the machine to sleep for a few minutes with the app open, then wake it — chat/report updates resume without needing to reload.
  • Verify that no errors appear in the JS console

Offline tests

The watchdog already early-returns while offline, so this change has no effect offline; behavior offline is unchanged.

QA Steps

// TODO: These must be filled out, or the issue title must include "[No QA]."

Same as tests

  • Verify that no errors appear in the JS console

PR Author Checklist

  • I linked the correct issue in the ### Fixed Issues section above
  • I wrote clear testing steps that cover the changes made in this PR
    • I added steps for local testing in the Tests section
    • I added steps for the expected offline behavior in the Offline steps section
    • I added steps for Staging and/or Production testing in the QA steps section
    • I added steps to cover failure scenarios (i.e. verify an input displays the correct error message if the entered data is not correct)
    • I turned off my network connection and tested it while offline to ensure it matches the expected behavior (i.e. verify the default avatar icon is displayed if app is offline)
    • I tested this PR with a High Traffic account against the staging or production API to ensure there are no regressions (e.g. long loading states that impact usability).
  • I included screenshots or videos for tests on all platforms
  • I ran the tests on all platforms & verified they passed on:
    • Android: Native
    • Android: mWeb Chrome
    • iOS: Native
    • iOS: mWeb Safari
    • MacOS: Chrome / Safari
  • I verified there are no console errors (if there's a console error not related to the PR, report it or open an issue for it to be fixed)
  • I followed proper code patterns (see Reviewing the code)
    • I verified that comments were added to code that is not self explanatory
    • I verified that any new or modified comments were clear, correct English, and explained "why" the code was doing something instead of only explaining "what" the code was doing.
    • I verified any copy / text that was added to the app is grammatically correct in English. It adheres to proper capitalization guidelines (note: only the first word of header/labels should be capitalized), and is either coming verbatim from figma or has been approved by marketing (in order to get marketing approval, ask the Bug Zero team member to add the Waiting for copy label to the issue)
  • If a new code pattern is added I verified it was agreed to be used by multiple Expensify engineers
  • I followed the guidelines as stated in the Review Guidelines
  • I tested other components that can be impacted by my changes (i.e. if the PR modifies a shared library or component like Avatar, I verified the components using Avatar are working as expected)
  • If a new CSS style is added I verified that:
    • A similar style doesn't already exist
    • The style can't be created with an existing StyleUtils function (i.e. StyleUtils.getBackgroundAndBorderStyle(theme.componentBG))
  • If new assets were added or existing ones were modified, I verified that:
    • The assets are optimized and compressed (for SVG files, run npm run compress-svg)
    • The assets load correctly across all supported platforms.
  • If the PR modifies code that runs when editing or sending messages, I tested and verified there is no unexpected behavior for all supported markdown - URLs, single line code, code blocks, quotes, headings, bold, strikethrough, and italic.
  • If the PR modifies a generic component, I tested and verified that those changes do not break usages of that component in the rest of the App (i.e. if a shared library or component like Avatar is modified, I verified that Avatar is working as expected in all cases)
  • If the PR modifies a component related to any of the existing Storybook stories, I tested and verified all stories for that component are still working as expected.
  • If the PR modifies a component or page that can be accessed by a direct deeplink, I verified that the code functions as expected when the deeplink is used - from a logged in and logged out account.
  • If the PR modifies the UI (e.g. new buttons, new UI components, changing the padding/spacing/sizing, moving components, etc) or modifies the form input styles:
    • I verified that all the inputs inside a form are aligned with each other.
    • I added Design label and/or tagged @Expensify/design so the design team can review the changes.
  • I added unit tests for any new feature or bug fix in this PR to help automatically prevent regressions in this user flow.
  • If the main branch was merged into this PR after a review, I tested again and verified the outcome was still expected according to the Test steps.

Screenshots/Videos

Android: Native

Non-UI change — verified via unit tests (tests/unit/PusherPingPongTest.ts), no visual changes.

Android: mWeb Chrome

Non-UI change — verified via unit tests (tests/unit/PusherPingPongTest.ts), no visual changes.

iOS: Native

Non-UI change — verified via unit tests (tests/unit/PusherPingPongTest.ts), no visual changes.

iOS: mWeb Safari

Non-UI change — verified via unit tests (tests/unit/PusherPingPongTest.ts), no visual changes.

MacOS: Chrome / Safari

Non-UI change — verified via unit tests (tests/unit/PusherPingPongTest.ts), no visual changes.

The one-shot latch disarmed the watchdog after a single reconnect, so a
socket that stayed dead (e.g. a network that throttles websockets) went
silently stale again. The PONG-clock reset now paces retries on its own
(~1 reconnect every 2 minutes), and the ping-send timestamp is left
alone so the fresh socket gets probed by the very next PING.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@adhorodyski

Copy link
Copy Markdown
Contributor Author

@codex review

@adhorodyski
adhorodyski marked this pull request as ready for review July 27, 2026 11:11
@adhorodyski
adhorodyski requested review from a team as code owners July 27, 2026 11:11
@melvin-bot
melvin-bot Bot requested review from JmillsExpensify and Krishna2323 and removed request for a team July 27, 2026 11:11
@melvin-bot

melvin-bot Bot commented Jul 27, 2026

Copy link
Copy Markdown

@Krishna2323 Please copy/paste the Reviewer Checklist from here into a new comment on this PR and complete it. If you have the K2 extension, you can simply click: [this button]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR improves the resiliency of the app’s Pusher “PING/PONG watchdog” by attempting to recover from stalled/dead websocket connections (e.g., after sleep/wake or websocket throttling) via Pusher.reconnect(), rather than only emitting a misleading “going offline” log.

Changes:

  • Update the Pusher PING/PONG watchdog to call Pusher.reconnect() when PONGs have been missing beyond the threshold.
  • Adjust watchdog log messaging to reflect the “socket presumed dead” behavior rather than “going offline”.
  • Add a new unit test covering the watchdog reconnect behavior and how a fresh PONG defers reconnects.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.

File Description
src/libs/actions/User.ts Changes watchdog behavior to reconnect Pusher when PONGs are missing past the threshold and updates related logging.
tests/unit/PusherPingPongTest.ts Adds unit coverage for watchdog reconnect timing and PONG-reset behavior.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/libs/actions/User.ts Outdated
Comment thread tests/unit/PusherPingPongTest.ts Outdated
Comment thread tests/unit/PusherPingPongTest.ts
adhorodyski and others added 3 commits July 27, 2026 16:45
…uite

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s deterministic

The previous approach reset the PONG clock on reconnect, which made the
next check tick a photo finish against the 60s threshold: timer jitter
(especially in throttled background tabs) could make the watchdog retry
every check tick (~1 minute) instead of every second one (~2 minutes),
doubling the ReconnectApp volume during an outage. Skipping exactly one
check has no time comparison to race, and keeping the PONG clock
untouched means logs now report the true age of the outage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

@JmillsExpensify JmillsExpensify left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No product review required.

Comment thread src/libs/actions/User.ts
@Krishna2323

Krishna2323 commented Jul 28, 2026 •

Copy link
Copy Markdown
Contributor

Reviewer Checklist

  • I have verified the author checklist is complete (all boxes are checked off).
  • I verified the correct issue is linked in the ### Fixed Issues section above
  • I verified testing steps are clear and they cover the changes made in this PR
    • I verified the steps for local testing are in the Tests section
    • I verified the steps for Staging and/or Production testing are in the QA steps section
    • I verified the steps cover any possible failure scenarios (i.e. verify an input displays the correct error message if the entered data is not correct)
    • I turned off my network connection and tested it while offline to ensure it matches the expected behavior (i.e. verify the default avatar icon is displayed if app is offline)
  • I checked that screenshots or videos are included for tests on all platforms
  • I included screenshots or videos for tests on all platforms
  • I verified that the composer does not automatically focus or open the keyboard on mobile unless explicitly intended. This includes checking that returning the app from the background does not unexpectedly open the keyboard.
  • I verified tests pass on all platforms & I tested again on:
    • Android: HybridApp
    • Android: mWeb Chrome
    • iOS: HybridApp
    • iOS: mWeb Safari
    • MacOS: Chrome / Safari
  • If there are any errors in the console that are unrelated to this PR, I either fixed them (preferred) or linked to where I reported them in Slack
  • I verified proper code patterns were followed (see Reviewing the code)
    • I verified that any callback methods that were added or modified are named for what the method does and never what callback they handle (i.e. toggleReport and not onIconClick).
    • I verified that comments were added to code that is not self explanatory
    • I verified that any new or modified comments were clear, correct English, and explained "why" the code was doing something instead of only explaining "what" the code was doing.
    • I verified any copy / text that was added to the app is grammatically correct in English. It adheres to proper capitalization guidelines (note: only the first word of header/labels should be capitalized), and is either coming verbatim from figma or has been approved by marketing (in order to get marketing approval, ask the Bug Zero team member to add the Waiting for copy label to the issue)
  • If a new code pattern is added I verified it was agreed to be used by multiple Expensify engineers
  • I verified that this PR follows the guidelines as stated in the Review Guidelines
  • I verified other components that can be impacted by these changes have been tested, and I retested again (i.e. if the PR modifies a shared library or component like Avatar, I verified the components using Avatar have been tested & I retested again)
  • If a new component is created I verified that:
    • A similar component doesn't exist in the codebase
    • All props are defined accurately and each prop has a /** comment above it */
    • The file is named correctly
    • The component has a clear name that is non-ambiguous and the purpose of the component can be inferred from the name alone
    • The only data being stored in the state is data necessary for rendering and nothing else
    • For Class Components, any internal methods passed to components event handlers are bound to this properly so there are no scoping issues (i.e. for onClick={this.submit} the method this.submit should be bound to this in the constructor)
    • Any internal methods bound to this are necessary to be bound (i.e. avoid this.submit = this.submit.bind(this); if this.submit is never passed to a component event handler like onClick)
    • All JSX used for rendering exists in the render method
    • The component has the minimum amount of code necessary for its purpose, and it is broken down into smaller components in order to separate concerns and functions
  • If any new file was added I verified that:
    • The file has a description of what it does and/or why is needed at the top of the file if the code is not self explanatory
  • If a new CSS style is added I verified that:
    • A similar style doesn't already exist
    • The style can't be created with an existing StyleUtils function (i.e. StyleUtils.getBackgroundAndBorderStyle(theme.componentBG)
  • If the PR modifies code that runs when editing or sending messages, I tested and verified there is no unexpected behavior for all supported markdown - URLs, single line code, code blocks, quotes, headings, bold, strikethrough, and italic.
  • If the PR modifies a generic component, I tested and verified that those changes do not break usages of that component in the rest of the App (i.e. if a shared library or component like Avatar is modified, I verified that Avatar is working as expected in all cases)
  • If the PR modifies a component related to any of the existing Storybook stories, I tested and verified all stories for that component are still working as expected.
  • If the PR modifies a component or page that can be accessed by a direct deeplink, I verified that the code functions as expected when the deeplink is used - from a logged in and logged out account.
  • If the PR modifies the UI (e.g. new buttons, new UI components, changing the padding/spacing/sizing, moving components, etc) or modifies the form input styles:
    • I verified that all the inputs inside a form are aligned with each other.
    • I added Design label and/or tagged @Expensify/design so the design team can review the changes.
  • For any bug fix or new feature in this PR, I verified that sufficient unit tests are included to prevent regressions in this flow.
  • If the main branch was merged into this PR after a review, I tested again and verified the outcome was still expected according to the Test steps.
  • I have checked off every checkbox in the PR reviewer checklist, including those that don't apply to this PR.

Screenshots/Videos

Android: HybridApp
Android: mWeb Chrome
iOS: HybridApp
iOS: mWeb Safari
MacOS: Chrome / Safari

@Krishna2323

This comment was marked as resolved.

@adhorodyski

Copy link
Copy Markdown
Contributor Author

@Krishna2323 does not repro for me

@Krishna2323 Krishna2323 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@melvin-bot
melvin-bot Bot requested a review from mountiny July 28, 2026 13:43
@mountiny mountiny changed the title fix: reconnect Pusher when a PONG is missed instead of only logging [noqa] fix: reconnect Pusher when a PONG is missed instead of only logging Jul 28, 2026
@mountiny
mountiny merged commit 1a17056 into Expensify:main Jul 28, 2026
35 of 39 checks passed
@OSBotify

Copy link
Copy Markdown
Contributor

✋ This PR was not deployed to staging yet because QA is ongoing. It will be automatically deployed to staging after the next production release.

@github-actions

Copy link
Copy Markdown
Contributor

🚧 mountiny has triggered a test Expensify/App build. You can view the workflow run here.

@OSBotify

Copy link
Copy Markdown
Contributor

🚀 Deployed to staging by https://github.com/mountiny in version: 9.4.46-0 🚀

platform result
🕸 web 🕸 success ✅
🤖 android 🤖 success ✅
🍎 iOS 🍎 success ✅

@MelvinBot

Copy link
Copy Markdown
Contributor

🤖 Help site review: no changes required

I reviewed the changes in this PR against Expensify's help site files under App/docs/articles and no documentation updates are needed, so I did not create a draft PR.

Why: This is a purely internal networking-resilience change. The Pusher PING/PONG watchdog now calls Pusher.reconnect() when a PONG has been missing for over 60 seconds (recovering the live connection after events like the machine sleeping or a throttled network), instead of only writing a misleading "going offline" log line. It also fixes that log text.

There is no user-facing change here — no new or modified feature, setting, tab, label, button, or workflow that a customer would read about. The effect is entirely behind the scenes: chat and report updates silently recover on their own instead of stalling until a manual reload. The help site documents product features and workflows, not the app's WebSocket connection internals, so there's nothing to add or update.

I also searched App/docs/articles for related terms (pusher, websocket, pong, reconnect) — the only matches are about reconnecting third-party integrations (e.g. Sage Intacct), which are unrelated to this connection-layer fix.

If you believe a customer-facing behavior did change and should be documented, let me know what to cover and I'll draft the article.

@OSBotify

Copy link
Copy Markdown
Contributor

🚀 Deployed to production by https://github.com/marcaaron in version: 9.4.46-10 🚀

platform result
🕸 web 🕸 success ✅
🤖 android 🤖 success ✅
🍎 iOS 🍎 success ✅

Bundle Size Analysis (Sentry):

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants