[ANCHOR-1251]: Permanent DoS of Stellar-RPC Payment Observer via Broken Fault-Recovery - #1983
Merged
Merged
Conversation
* fix: rpc observer restart executor service after shutdown * add: generic runtime exception catch in restart logic * refactor: checkstatus to handle unexpected internal exceptions * add: e2e test for rpc observer recovery after outages * add: unit tests for observer executor and error handling
Contributor
There was a problem hiding this comment.
Pull request overview
Hardens payment observer fault recovery to prevent polling from permanently stopping after transient failures.
Changes:
- Recreates terminated Stellar RPC polling executors during restart.
- Guards restart and supervisor paths against runtime exceptions.
- Adds unit and lifecycle recovery tests.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
AbstractPaymentObserver.java |
Adds supervisor exception handling. |
StellarRpcPaymentObserver.java |
Recreates the polling executor after shutdown. |
StellarRpcPaymentObserverTest.kt |
Adds restart regression tests. |
StellarRpcPaymentObserverRecoveryE2ETest.kt |
Adds transient-outage recovery coverage. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
* refactor stellar rpc observer recovery e2e test wait logic * add assertions for executor replacement and shutdown in recovery test * update stellar rpc payment observer restart internal test mocks and verification
JiahuiWho
approved these changes
Jul 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
StellarRpcPaymentObserverreuses a single-shotScheduledExecutorServiceacross restarts.shutdownInternal()callsexecutorService.shutdownNow(), which terminates it permanently, but the field is never reassigned. The first time any transient fault (RPC disconnect, a silence timeout, a DB hiccup, an event-publisher error) triggersAbstractPaymentObserver.restartInternal(),startInternal()tries to reschedule onto that already-terminated executor and throwsRejectedExecutionException. That exception was only ever caught asTransactionExceptioninrestartInternal(), andcheckStatus()— the body of thestatusWatcher'sscheduleWithFixedDelaytask — had no guard around it at all, so the exception escaped both. PerScheduledExecutorService's contract, an uncaught exception in a fixed-delay task permanently suppresses every future execution of that task. The result: one ordinary transient fault permanently and silently halts all on-chain payment detection until a manual process restart.Fixing only the executor reuse isn't sufficient on its own.
restartInternal()andcheckStatus()are shared by bothStellarRpcPaymentObserverandHorizonPaymentObserver— Horizon isn't vulnerable to this specific trigger only because it happens to build a freshSSEStreamon every restart, not because the supervisor itself is hardened. A future change to Horizon's own lifecycle could reintroduce this same class of failure there. The supervisor loop itself needed to become crash-proof, independent of which subclass or which fault caused the failure.Changes
StellarRpcPaymentObserver.startInternal: recreatesexecutorServicewhen it's null or already shut down, instead of always scheduling onto the same single-shot instance created at construction. This is the actual root-cause fix — restarts no longer hit a terminated executor.AbstractPaymentObserver.restartInternal: broadened the catch fromTransactionExceptiononly to also catchRuntimeException, settingSTREAM_ERROR. This status choice is deliberate:restartInternal()is only ever called from withincheckStatusInternal()'s ownSTREAM_ERROR/SILENCE_ERROR/PUBLISHER_ERROR/DATABASE_ERRORbranches, each of which has its own bounded backoff-then-shutdown logic (streamBackoffTimer.isTimerMaxed(),silenceTimeoutCount, etc).STREAM_ERRORis a no-op when set from theSTREAM_ERRORbranch itself and a rejected, harmless transition from the other three (perObserverStatus.stateTransition, which only allowsSTREAM_ERRORfromRUNNING) — either way, control returns with the original error status intact, so each branch's own retry count/timer still advances normally on the next tick instead of being short-circuited.AbstractPaymentObserver.checkStatus: split into a thin wrapper around the renamedcheckStatusInternal(), catching anyRuntimeExceptionas a last-resort backstop so the supervisor task itself can never die. Falls back toNEEDS_SHUTDOWNrather thanSTREAM_ERRORhere, sinceNEEDS_SHUTDOWNis the only status reachable from every error state in the transition table —STREAM_ERRORwould be silently rejected in most of them, leaving the failure invisible instead of converging to a clean, observable stop.StellarRpcPaymentObserverTest.kt: added regression tests — executor recreation after shutdown,restartInternalgenuinely resuming polling (verified viasorobanServer.getEventsactually being invoked again),restartInternalswallowing the exactRejectedExecutionExceptionthe defect throws, andcheckStatus's backstop converging toNEEDS_SHUTDOWNwhen something unexpected throws.StellarRpcPaymentObserverRecoveryE2ETest.kt(new): drives the observer's realstart()lifecycle — actualexecutorService/silenceWatcher/statusWatcherthreads on real wall-clock timing, only theSorobanServernetwork boundary mocked — through an induced transient outage pastsilence_timeoutand back, asserting polling genuinely resumes and health returns toGREEN.Acceptance Criteria
restartInternal(), the observer resumes polling and its health check returns toGREEN, with no process restart required.StellarRpcPaymentObserver.startInternal()never throwsRejectedExecutionExceptionaftershutdownInternal()has run.RuntimeExceptionanywhere incheckStatus()'s switch body never permanently stops the status watcher's scheduled task.streamBackoffTimer.isTimerMaxed(),silenceTimeoutCountvssilenceTimeoutRetries, etc.) is unchanged for restarts that succeed normally.Context
HackerOne #3857031
Testing
./gradlew :platform:test --tests "org.stellar.anchor.platform.observer.stellar.StellarRpcPaymentObserverTest"./gradlew :platform:test --tests "org.stellar.anchor.platform.observer.stellar.StellarRpcPaymentObserverRecoveryE2ETest"Documentation
N/A
Known limitations
N/A