Skip to content

Update embeddings server: return HTTP 400 for invalid embedding requests - #29060

Merged
ngxson merged 1 commit into
ggml-org:masterfrom
SamMalayek:fix/embedding-invalid-request
Oct 1, 2026
Merged

ngxson merged 1 commit into
ggml-org:masterfrom
SamMalayek:fix/embedding-invalid-request

Conversation

@SamMalayek

@SamMalayek SamMalayek commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Overview

  • Returns HTTP 400 invalid_request_error instead of HTTP 500 for malformed embedding requests that fail during request parsing or validation.
  • Throws std::invalid_argument instead of std::runtime_error for clearly client-caused tokenizer failures.
  • Catches common_json_error during embedding request preparation.
  • Adds test coverage for invalid /v1/embeddings requests and /embeddings cases.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES -- AI was used to assist with investigation, implementation, and testing. I manually reviewed the changes and validated them.

@SamMalayek
SamMalayek marked this pull request as ready for review September 18, 2026 00:36
@SamMalayek
SamMalayek requested a review from a team as a code owner September 18, 2026 00:36
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown

Hi @SamMalayek, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@SamMalayek

SamMalayek commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Hi @SamMalayek, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

This is a false positive (it says Contributor in my post headings -- example merge: #16541).

And as with all my PRs after I became familiar with llama.cpp & OSS, this is a clean and optimal change that is simply needed.

@SamMalayek

SamMalayek commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Could a maintainer rerun the failed + cancelled CI jobs, please?

We've got 503, cancelled, and similar failures that appear transient. Furthermore, the 3x self-hosted jobs have been queued for hours (update: are now auto-cancelled). Is there an issue with runner availability?

Thanks.

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

CI restarted for failed jobs

@SamMalayek

SamMalayek commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

@ServeurpersoCom : Server / ubuntu still appears to be the original Hugging Face 503 failure and was never rerun. The 3x Intel self-hosted jobs also timed out after 24h without getting runners. Could you rerun Server / ubuntu? Also, is anything needed for the unavailable Intel runners?

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Hi, returning 400 is legit per the OpenAI spec, since the SDK retries every 5xx and a malformed request ends up sent three times.

This needs a rebase on master, and at first read I think the fix can be much more targeted: common_json_error is the only reason for the try block, so could you catch it once in ex_wrapper as invalid_request_error? That keeps the C++ side to that catch plus the two invalid_argument changes, and covers every route that parses the body.

The test can also be lighter, a single parametrize over the invalid bodies asserting the 400, like the rest of test_embedding.py.

@SamMalayek

SamMalayek commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

Hi, returning 400 is legit per the OpenAI spec, since the SDK retries every 5xx and a malformed request ends up sent three times.

This needs a rebase on master, and at first read I think the fix can be much more targeted: common_json_error is the only reason for the try block, so could you catch it once in ex_wrapper as invalid_request_error? That keeps the C++ side to that catch plus the two invalid_argument changes, and covers every route that parses the body.

The test can also be lighter, a single parametrize over the invalid bodies asserting the 400, like the rest of test_embedding.py.

Thanks for the guidance. I’ve implemented the ex_wrapper approach and simplified the tests (removing edge case testing, but following guidance and convention). One implication of this ex_wrapper change is that an uncaught common_json_error arising during internal/response-side JSON processing could theoretically & incorrectly be classified as a 400 rather than a 500 (unlikely if high caution is utilized by all devs). I’m proceeding with this approach as suggested.

ex_wrapper can’t reliably distinguish client vs internal common_json_error without broader changes, so again, I’m following the centralized 400 handling you suggested. My original code had the same issue, but much narrower scope and blast radius (also better position to iterate towards a narrow 100% accurate solution, but you could argue that the best solution is to broaden the scope of ex_wrapper... might come back to this later). TLDR: My original solution was more surgical and better in the short term but this new one is more broad (since ex_wrapper is used by so many endpoints) and better positioned for the long term. Both solutions are better than leaving the code as it was.

@SamMalayek
SamMalayek force-pushed the fix/embedding-invalid-request branch from 72b074a to b139644 Compare September 22, 2026 21:18
@ServeurpersoCom

ServeurpersoCom commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Closing and reopening to retrigger the CI, this run still uses the cmake-pkg workflow from before #29299 which lands on runners that only accept container jobs, no action needed on your side.

Closing and reopening makes the CI redo the clone rebased on current master on its side, so as long as there is no conflict there is no need to rebase.

@SamMalayek

Copy link
Copy Markdown
Contributor Author

Closing and reopening to retrigger the CI, this run still uses the cmake-pkg workflow from before #29299 which lands on runners that only accept container jobs, no action needed on your side.

Closing and reopening makes the CI redo the clone rebased on current master on its side, so as long as there is no conflict there is no need to rebase.

@ServeurpersoCom All 17 CI checks are passing now. Could you take another look at the updated PR?

@ServeurpersoCom ServeurpersoCom left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks for reworking it around ex_wrapper: malformed bodies now get a proper 400 instead of a 500, so OpenAI-compatible clients stop retrying them.

@ngxson
ngxson merged commit d775ebf into ggml-org:master Oct 1, 2026
37 of 40 checks passed
thom-dev-fr added a commit to thom-dev-fr/llama.cpp that referenced this pull request Oct 2, 2026
Upstream ggml-org#29060 maps common_json_error to HTTP 400. Requests prepared by
the engine reported them as preparation_failed (500), which failed the new
test_embedding_invalid_request case for a non-string encoding_format.
Malformed bodies now return 400 as upstream: update test_sleep accordingly.
thom-dev-fr added a commit to thom-dev-fr/llama.cpp that referenced this pull request Oct 2, 2026
Upstream ggml-org#29060 maps common_json_error to HTTP 400. Requests prepared by
the engine reported them as preparation_failed (500), which failed the new
test_embedding_invalid_request case for a non-string encoding_format.
Malformed bodies now return 400 as upstream: update test_sleep accordingly.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants