Skip to content

BUG: Frontend aborts long operations after 10s (global axios timeout) - reclustering, memory generation & search fail on large libraries #1345

Description

@VanshajPoonia

Is there an existing issue for this?

  • I have searched the existing issues

What happened?

The shared axios client applies a hard 10-second timeout to every backend request (frontend/src/api/axiosConfig.ts:7 → timeout: 10000; the sync client at :15 too), with no per-request override anywhere in frontend/src (verified by grep).

Several backend endpoints are synchronous and scale with library size, so on a realistic library they run past 10s. The client aborts with a timeout/network error and shows a failure to the user — even though the backend keeps running to completion — which can leave the UI and DB inconsistent.

Affected calls (all use the 10s apiClient):

  • triggerGlobalReclustering → POST /face-clusters/global-recluster (frontend/src/api/api-functions/face_clusters.ts:78). Backend runs full DBSCAN over all embeddings + regenerates a face image per cluster.
  • generateMemories → POST /api/memories/generate (frontend/src/api/api-functions/memories.ts:73). Clusters every image with location data.
  • getTimeline / getLocations, fetchSearchedFaces, fetchMultiPersonSearch — same client, can exceed 10s on large libraries.

Steps to reproduce:

  1. Add folders with a few thousand photos; enable AI tagging and let faces populate.
  2. Trigger Recluster faces (or Memories → Generate).
  3. Request fails after ~10s in the UI; backend logs show it still working.

Expected: long operations either run without an artificial 10s cap, or run asynchronously with progress.

Proposed fix:

  • Short term: remove the global 10s default (or raise it substantially) and set conservative per-call timeouts only where they make sense.
  • Better: convert global-recluster and memories/generate to an async job + polling/SSE, reusing the existing model-download SSE pattern in backend/app/routes/models.py.

Not a duplicate: searched all issues (title + body) and all PRs for timeout/axios; only matches are SQLite busy_timeout work (PR #866/#982), unrelated to the HTTP client.

Record

  • I agree to follow this project's Code of Conduct

Activity

  1. github-actions commented on Jun 27, 2026

    @github-actions
    Contributor

    👋 Thanks for opening this issue! The maintainers will review it shortly. Till then, do not open any PR. This is a strict rule we like to follow around here :)

  2. VanshajPoonia commented on Jun 27, 2026

    @VanshajPoonia
    ContributorAuthor

    @rohan-pandeyy Can I work on this issue?

  3. rohan-pandeyy commented on Jul 3, 2026

    @rohan-pandeyy
    Member

    @VanshajPoonia can you confirm whether issue #1355 overlaps with your work within this issue?

  4. VanshajPoonia commented on Jul 3, 2026

    @VanshajPoonia
    ContributorAuthor

    @rohan-pandeyy Partial overlap, on one endpoint only:

    POST /face-clusters/global-recluster - the one @aliviahossain flags in #1355 as "fully synchronous, blocks the request thread... needs task ID + polling treatment" - is exactly what PR #1349 (currently open) already implements: it returns a task_id immediately and adds GET /face-clusters/global-recluster/{task_id} for polling status.

    The other five endpoints she lists (add-folder, enable-ai-tagging, sync-folder, face-search, memories/generate) are untouched by #1349 and still need the work she describes - face-search only got a longer timeout, not an async conversion, and memories/generate was deliberately left alone since it's owned by a separate GSoC effort.

    So: no conflict if #1355 is scoped to those five, but worth pointing her at #1349 so she doesn't duplicate the global-recluster piece.

  5. VanshajPoonia commented on Jul 20, 2026

    @VanshajPoonia
    ContributorAuthor

    Hey @rohan-pandeyy, following up on your question about the #1355 overlap:
    that issue has since been closed as out of scope, so the overlap concern is
    moot and this PR is unblocked.

    Current state: rebased cleanly on main (no conflicts), all CI green,
    CodeRabbit approved, no unresolved threads. The underlying bug is still
    live, axiosConfig.ts on main still has the 10s cap on both clients.

    Could you take a look when you get a chance?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

backendbugSomething isn't working

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions