Skip to content

[Bug]: a local Windows ACL hardening failure is returned to the client as 401 authentication_error #1296

Description

@brunoflma

Client or integration

Codex CLI

Area

Authentication and account pool

Summary

A local filesystem problem — Windows ACL hardening of a secret path exceeding its budget — reaches the client as 401 authentication_error. The user sees what looks like an expired or rejected credential, and the actual message can be the literal ACL hardening skipped - previous attempt timed out.

Two properties combine to produce this:

  1. hardenSecretPath(..., { required: true }) throws on a genuine timeout, and the timeout memo timedOutPaths is process-global (src/lib/windows-secret-acl.ts:47). Once a path times out, subsequent writes to it keep failing for the life of the process — one authorized recovery attempt aside.
  2. src/server/responses/core.ts passes the raw exception message straight into an auth error in three branches (:880, :883, :886):
return { ok: false, response: formatErrorResponse(401, "authentication_error", err.message) };

When the poisoned path is auth.json, an OAuth refresh cannot persist, and the resulting failure is reported as a credential error rather than as a local storage/permission error. src/server/management-auth.ts is more careful on its own surface — it says management token directory ACL hardening did not complete (:76, :95) — but the data-plane path has no such classification.

Expected: an ACL/filesystem failure should surface with its own classification (a 5xx, or at minimum an error type that is not authentication_error), naming the local cause. Whatever the status code, it should not be indistinguishable from "your token is bad", because the two have opposite remediations — one is chmod/ACL/disk, the other is re-login.

This is a diagnosability defect rather than a data-loss one, but it cost me a long misdiagnosis: I went looking at provider credentials while the actual cause was a directory tree under ~/.opencodex.

Reproduction

Honest status first: I no longer have a live reproduction, because I fixed the local trigger. I am reporting it because the mechanism is unchanged in 2.11.0 and the misclassification is structural, not incidental. The template invites this case, so here is exactly what I observed and under which conditions.

What produced it originally, on 2.10.x:

  1. ~/.opencodex had grown to ~94,700 descendants (git worktrees and backups under it).
  2. hardenSecretDir applies /grant:r <principal>:(OI)(CI)(F) followed by /inheritance:r. (OI)(CI) makes Windows propagate the ACE to every descendant, so the write cost is linear in the number of descendants — measured on this host: 88 ms at 10 files, 9,438 ms at 20,000, extrapolating to ~45 s at 94,715.
  3. That exceeded the harden budget, hardenSecretPath(required: true) threw, and the path entered the global timeout memo.
  4. POST /v1/responses then returned 401 authentication_error carrying the literal text ACL hardening skipped - previous attempt timed out. The service log had 319 occurrences of the ACL failure up to the day I fixed it.
  5. The dashboard surface failed differently and more usefully, with GET /api/logs → 503 {"error":"management API unavailable"} and reason management token directory ACL hardening did not complete.

To recreate the condition synthetically on Windows, without waiting for a tree to grow:

  1. Place a large subtree under ~/.opencodex (tens of thousands of files) with ACL inheritance enabled.
  2. Lower the envelope: OPENCODEX_ACL_TIMEOUT_MS=1000.
  3. Restart the proxy and send POST /v1/responses for any routed provider.
  4. Observe a 401 authentication_error whose message describes ACL hardening rather than a credential.

Reading the code path alone is enough to see the classification issue: windows-secret-acl.ts throws a sanitized ACL error, and responses/core.ts:880-886 re-labels whatever message arrives as authentication_error.

Note on 2.11.0: #1156 raised HARDEN_DEADLINE_DEFAULT_MS from 5,000 to 30,000 (src/lib/windows-secret-acl.ts:253), which widens the margin — but it does not change the classification, and the process-global memo is unchanged.

Version

2.11.0

Operating system

Windows 11 Home Single Language 26H2 (build 10.0.26200)

Provider and model

Not provider-specific — the failure is in credential persistence, before provider selection.

Logs or error output

# data plane — the misleading one
POST /v1/responses -> 401 authentication_error
  message: "ACL hardening skipped - previous attempt timed out"

# management plane — same root cause, correctly classified
GET /api/logs -> 503 {"error":"management API unavailable"}
  reason: "management token directory ACL hardening did not complete"

# measured propagation cost of `/grant:r (OI)(CI)` on this host
10 files      ->     88 ms
20,000 files  ->  9,438 ms
94,715 files  -> ~45 s (extrapolated; above the 2.10.x budget)

Screenshots and supporting files

Not attached.

Redacted configuration

{
  "note": "Not configuration-dependent. Relevant environment only:",
  "OPENCODEX_ACL_TIMEOUT_MS": "5000",
  "platform": "win32"
}

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

account-poolOAuth, credentials, Codex pool, quota, failover, plansbugSomething isn't workingplatformOS/service/tray/ACL (Windows-heavy, not Windows-only)proxyHTTP proxy, routing, reverse-proxy / management auth

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions