Skip to content

[Bug]: Windows startup can run the 20 s schtasks fallback twice before listen #2923

Description

@Ingwannu

Confirmed current-dev regression

Current dev 64b994c includes #2918. On Windows, a localized nonzero targeted schtasks query falls back to the full task listing with a 20 second ceiling.

startServer currently calls inspectStartupOwnership before startup cache invalidation at src/server/index.ts around line 609 and again before native-main lifecycle preparation around line 788. Both calls reach the same synchronous service-manager probe.

On the reported zh-CN host, the full listing measured about 12.3 seconds for 401 tasks. With no installed OpenCodex task, one proxy start can therefore spend about 25 seconds in two identical listings before Bun.serve starts, with a 40 second worst-case ceiling.

The second ownership check is intentional. It prevents startup work between the cache decision and native lifecycle preparation from widening the service-install race. The fix must not simply delete final revalidation, lower the listing timeout back to 2 seconds, add guessed absence, or convert an unknown result into absent.

Required scope

  • Keep each final ownership decision fail-closed.
  • Keep the targeted scheduled-task query bounded and revalidated.
  • Avoid repeating the expensive full-machine enumeration during one startup when the targeted recheck still provides the race-sensitive evidence.
  • Do not broaden this into service installation or GUI changes.

Regression coverage

Add an injected Windows startup probe that:

  1. makes the targeted query return a localized unreadable nonzero result,
  2. makes the full listing take the fallback path,
  3. proves one start does not run the full listing twice,
  4. proves the targeted ownership recheck still occurs, and
  5. proves a task that becomes present between checks is not reported absent.

Original review evidence: #2920 (comment)

Activity

  1. added
    platformOS/service/tray/ACL (Windows-heavy, not Windows-only)
    on Aug 29, 2026
  2. lidge-jun commented on Aug 29, 2026

    @lidge-jun
    Owner

    리뷰 · 우선순위 70 / 80

    이 이슈는 Windows 에서 프록시가 듣기 전에 같은 비싼 일을 두 번 하는 버그다. 지금 dev HEAD 는 64b994c0f 이다. #2918 이 이미 들어가 있다. 그 고침은 zh/ja 호스트에서 schtasks 출력을 제대로 읽고, 전체 작업 목록만 SERVICE_PROBE_LISTING_TIMEOUT_MS = 20_000 으로 길게 기다린다. 대상 조회는 예전처럼 2초다. 문제는 그 긴 전체 목록이 한 번의 시작 안에서 두 번 돌 수 있다는 점이다.

    코드 위치는 분명하다. src/server/index.ts 의 startServer 가 inspectStartupOwnership 을 609행 근처에서 한 번 부른다. 캐시를 지울지 말지 정하기 위해서다. 그다음 788행 근처에서 또 부른다. 네이티브 메인 생명주기를 준비하기 직전이다. 두 번째 호출은 일부러 있다. 캐시 결정과 생명주기 준비 사이에 서비스가 설치되면, 예전 결과만 믿으면 레이스가 커지기 때문이다. 815행의 재시도 경로도 같은 헬퍼를 쓴다. 헬퍼 자체는 498행에 있고, 결국 inspectNativeCodexOwnership → Windows 면 service-manager-probe 의 schtasks 경로로 간다.

    재현 조건도 좁다. 대상 /query /tn ... 가 로케일 때문에 nonzero 로 실패하고, 그때만 전체 목록으로 넘어간다. 본문이 잰 zh-CN 호스트는 작업 401개에 목록이 약 12.3초였다. OpenCodex 작업이 없으면 한 번 시작에 목록이 두 번 돌아가 약 25초가 듣기 전에 쓰이고, 최악은 40초 천장이다. #2918 이 목록을 살린 대가인데, 같은 시작 안에서 두 번 내는 건 새 회귀다.

    고치면 안 되는 것도 이슈가 잘 적어 두었다. 두 번째 소유권 재확인을 지우면 안 된다. 목록 타임아웃을 다시 2초로 내리면 #2914 가 돌아온다. 추측으로 없다 하지 말고, unknown 을 absent 로 바꾸지 말라. 필요한 건 대상 재확인은 유지한 채, 한 시작 안에서 값비싼 전체 목록만 한 번으로 묶는 것이다.

    라인 - src/server/index.ts 609행과 788행 - 같은 inspectStartupOwnership 을 한 시작에 두 번 부른다. 둘 다 대상 조회가 실패하면 전체 목록(최대 20초)으로 떨어진다.
    라인 - src/service-manager-probe.ts 46행 SERVICE_PROBE_LISTING_TIMEOUT_MS - 목록만 20초다. 대상 조회는 33행의 2초를 유지한다. 타임아웃을 줄이는 해결은 아니다.
    경로/심볼 - inspectStartupOwnership / inspectNativeCodexOwnership - 주입 가능한 Windows 시작 프로브로 목록 횟수·대상 재확인·중간에 생긴 작업을 absent 로 안 보고를 증명해야 한다. 본문이 적은 다섯 가지 회귀 조건이 합격선이다.

    메인테이너의 판단이 필요한 지점

    • 시작 한 번 동안 전체 목록 결과를 메모이즈할지, 아니면 대상 조회만 재실행하고 목록은 캐시할지.
    • 메모이즈 수명을 프로세스 시작 구간으로만 둘지, 프로브 레이어에 짧은 TTL을 둘지.
    • #2918 직후 플랫폼 핫픽스로 바로 받을지, 다른 Windows 서비스 PR과 묶을지.

    너의 추천
    우선순위 높게 받아라. #2918 이 살린 목록이 시작마다 두 번 나가면 Windows 체감이 다시 나빠진다. 범위는 서비스 설치·GUI 로 넓히지 말고, 시작 한 번의 목록 중복만 없애라. 본문의 주입 프로브 다섯 조건을 테스트로 잠근 뒤 dev 에 넣어라. 관련 증거는 #2920 리뷰 스레드에 있다.

    이 댓글은 grok-bot이 작성했습니다

  3. Ingwannu commented on Aug 29, 2026

    @Ingwannu
    OwnerAuthor

    Confirmed and fixed in #2928. The patch keeps both targeted Task Scheduler ownership queries, shares only the expensive full listing when the fresh targeted result is byte-for-byte unchanged, and leaves later runtime retries uncached. A changed query result or listing failure remains fail-closed, and a task appearing between the two startup checks is not reported absent.\n\nVerification: 112 focused tests passed on the exact PR head, typecheck passed, and the full suite completed with 16,144 pass / 16 skip / 0 fail before the final non-overlapping dev rebase. Owner review and exact-head CI are requested; I am not merging until those gates are satisfied.

  4. lidge-jun commented on Aug 29, 2026

    @lidge-jun
    Owner

    Fixed on dev in the merge below (#2928, by @Ingwannu).

    Your diagnosis was exactly right, including which fixes would be wrong — deleting the revalidation, lowering the budget back to 2 s, or converting unknown into absent were all genuinely tempting and all wrong.

    The listing is now reused only while the targeted /query /tn result is byte-for-byte unchanged. That is stronger than scoping it to "one startup," which is what I had written in #2927 before closing it in favor of this: binding the absence proof to the evidence that produced it means changed targeted evidence invalidates the listing immediately, rather than relying on ordering within a time window. Both ownership checks keep their fresh targeted query, so the install race the second check exists to close is untouched, and runtime retries deliberately omit the memo.

    One gap I found in review and fixed on the branch: the first version cached a stalled listing too, so a single 20 s timeout left ownership unprovable for the rest of the startup — refusing exactly the write #2914 exists to allow. Only a successful listing is retained now, and there is a regression test where the targeted stderr is byte-identical across both passes so nothing but the stall-handling can force the retry.

    All five coverage points you asked for are in tests/codex-service-manager-probe-hardening.test.ts, and each was mutation-proven rather than assumed green.

    What was not verified: no real schtasks ran — the implementer and I are both on macOS — so the call-count and caching behavior are covered by injected probes and the ~25 s wall-clock claim rests on your measurements.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingplatformOS/service/tray/ACL (Windows-heavy, not Windows-only)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions