Skip to content

Fix reserved characters in GitHub content file paths - #2198

Merged
martindurant merged 1 commit into
fsspec:masterfrom
rioyu123:codex/github-content-path-encoding
Sep 28, 2026
Merged

martindurant merged 1 commit into
fsspec:masterfrom
rioyu123:codex/github-content-path-encoding

Conversation

@rioyu123

Copy link
Copy Markdown
Contributor

GithubFileSystem inserted file paths into GitHub contents API URLs without encoding them. A # in a filename was treated as a URL fragment, ? started a query string, and a literal %23 or %2F was decoded into a different character. As a result, opening or removing such a file could fail with a 404 or act on a different path.

This change percent-encodes the path when building the read, SHA-lookup and delete URLs. Directory separators are left as they are, so nested paths work as before, and names returned by ls() can now be passed straight to open() and rm().

Branch/ref handling is unchanged. Callers who previously pre-encoded file paths to work around this should now pass literal names.

Tests: a new offline test module patches requests.Session.send, so Requests still prepares the URLs for real but no network calls are made. It covers #, ?, literal %23/%2F, spaces, Unicode and plain names, and deletion with and without a cached SHA. The tests check the exact encoded URL path and the request sequence. The 10 reserved-character cases fail on the base commit, and all 13 pass with the fix. The related spec/core/utils selection passes (306 passed, 129 skipped, 1 xfailed), with one Windows HTTP server teardown warning. Changed-file Ruff 0.14.3 and whitespace checks pass. No live GitHub requests were made by the new tests.

@martindurant
martindurant merged commit 49825f3 into fsspec:master Sep 28, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants