Skip to content

feat: delete-failed-jobs-flag - #838

Closed
Powfu2 wants to merge 1 commit into
timgit:masterfrom
Powfu2:master
Closed

Powfu2 wants to merge 1 commit into
timgit:masterfrom
Powfu2:master

Conversation

@Powfu2

@Powfu2 Powfu2 commented Jul 7, 2026

Copy link
Copy Markdown

Adds a new boolean configuration flag (e.g. deleteFailedJobs) that controls whether failed jobs are included in the supervisor's retention cleanup.

When the flag is true (default), behavior is unchanged — both completed and failed jobs are deleted per the retention policy.

When the flag is false, the supervisor only deletes completed jobs, and failed jobs are kept in the table for debugging, auditing, and manual reprocessing.

@blue2cat

blue2cat commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

Thanks for the PR! I’m not the maintainer, but I did want to note that this use case may already be covered via a dead letter queue.

For keeping failed jobs around for debugging, or manual reprocessing, DLQs already preserve the payload and failure output while not changing the source queue’s retention behavior.

If the goal here is to prune successful jobs while retaining failed source rows, maybe that would fit better as a per-queue retention option instead of a constructor flag?

@Powfu2

Powfu2 commented Jul 8, 2026

Copy link
Copy Markdown
Author

Thanks for the thoughtful review @blue2cat

You're right that DLQs cover a good chunk of this — with the source* fields
(sourceName, sourceId, etc.) debugging from a DLQ is quite workable.

What I'm after is slightly different though: keeping the original failed rows
in the source queue for auditing and history, without duplicating them into a
second queue that has its own retention lifecycle and requires extra setup per queue.

That said, I agree the per-queue suggestion is the better shape — especially since
retention config already lives at the queue level (deleteAfterSeconds,
retentionSeconds). I'll rework this PR as a queue option instead of a constructor
flag. My current thinking:

  • deleteFailedAfterSeconds?: number on createQueue() / updateQueue()
  • Defaults to the existing deleteAfterSeconds behavior (no breaking change)
  • A way to opt out of deletion entirely for failed jobs, so they're retained
    indefinitely for auditing/manual reprocessing

This would be symmetrical with the existing API and lets users prune completed
jobs aggressively while keeping failed rows around.

Does that sound reasonable? Happy to adjust before pushing the rework.

@blue2cat

Copy link
Copy Markdown
Contributor

@Powfu2, per-queue options make more sense to me, but I'd be curious to get @timgit's thoughts on the addition.

@timgit

timgit commented Jul 24, 2026

Copy link
Copy Markdown
Owner

I'll rework this PR as a queue option instead of a constructor flag.

I agree with switching away from a new constructor option, but I'm hesitant to add another option that encourages manual intervention to prevent a maintenance issue, such as a table that can never remove its data.

keeping the original failed rows in the source queue for auditing and history, without duplicating them into a
second queue that has its own retention lifecycle and requires extra setup per queue.

What if you were to create a single DLQ that catches failures from all of your queues? It would only require a one-time retention configuration in that case. There are probably plenty of reasons to not do this, but it may work if your failure rates are under control.

@timgit

timgit commented Sep 25, 2026

Copy link
Copy Markdown
Owner

Another option you can do is use deleteJob() in your handlers to drop all of the successfully processed jobs as they're completed. If it's too soon, create a new queue 'delete-job' to defer it by a bit.

I'll go ahead and close this now as it's a bit stale.

@timgit timgit closed this Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants