Repository navigation
server: reject partial media truncation - #24076
Conversation
|
can you provide an example request (maybe as a python script) that can fully reproduce the issue? |
There was a problem hiding this comment.
remove this, we don't do server test this way
There was a problem hiding this comment.
Done in the current branch. The tests/test-server-tokens.cpp coverage was removed, and the request-level regression now lives in tools/server/tests/unit/test_vision_api.py following the existing server test structure.
Current checks only show labeler passing; I do not see a fresh failing server-test run on the latest head yet.
bba9d3e to
ce99f9e
Compare
|
Thanks, removed A request-level repro needs a multimodal server and a context shift where import base64
import requests
base = "http://127.0.0.1:8080"
props = requests.get(f"{base}/props", timeout=10).json()
marker = props["media_marker"]
# 1x1 PNG. Any valid image works; this just keeps the repro small.
img = (
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk"
"+A8AAQUBAScY42YAAAAASUVORK5CYII="
)
prompt = {
"prompt": "alpha beta gamma delta epsilon " + marker + " zeta eta theta iota kappa lambda mu",
"multimodal_data": [img],
}
payload = {
"prompt": prompt,
"cache_prompt": True,
"n_keep": 6,
"n_discard": 1,
"n_predict": 64,
"temperature": 0,
}
# With a small enough --ctx-size, this forces the shift path. The exact n_keep
# can be adjusted by +/- a few tokens depending on the tokenizer/template; the
# failing condition is that n_keep is between the first and last token of the
# media chunk.
for _ in range(2):
r = requests.post(f"{base}/completion", json=payload, timeout=120)
print(r.status_code, r.text[:500])The old path checked |
|
can you add one single test case to |
ce99f9e to
00c3547
Compare
|
Added the requested request-level coverage in The new test sends a multimodal Local validation: I also tried to run the focused pytest locally, but this Windows MinGW build is currently blocked before producing The first build attempt needed |
|
the test on CI didn't pass |
00c3547 to
97cef89
Compare
|
Updated in 97cef893. The CI failure was caused by the added request-level test hitting context-size validation before it reached the media-boundary path. I changed the test to keep the request below the context limit, prime the slot cache with a successful multimodal completion, then repeat the same cached prompt so the cache trim lands inside the media chunk and returns Local validation: I also retried the focused pytest locally after installing the server test requirements. It is still blocked by this Windows environment: rebuilding |
|
I rechecked the current head ( For the earlier failing run on So I am leaving the code unchanged for now; this branch needs the workflows approved/rerun before there is a fresh signal on the current test. |
|
Rechecked the current state: the only requested-change thread (removing the old C++ server test) is now outdated, because the branch moved the coverage into The current blocker appears to be |
|
sorry for the delay, I think this now needs a rebase |
97cef89 to
953acde
Compare
|
Rebased onto current The earlier review point (dropping the old C++ |
|
Re-verified on current master: the |
|
Hi, I tried the new test locally: tinygemma3 uses SWA so the server reprocesses the whole prompt (n_past = 0) and keep_first never cuts, the second request returns 200 with and without the fix so the assertion fails, the fix itself looks right though. |
|
not sure why the CI doesn't trigger |
|
merging this when the CI passes |
Drop the mtmd test helper change, which no longer builds since clip_image_f32_batch stores its entries by value, and drop the vision test: no test fixture reaches a cut between two adjacent media chunks with a reused cache (tinygemma3 uses SWA and wraps images in text tokens, tinyopenjev and small-test are recurrent), so the test passed or failed independently of the fix.
953acde to
ca800b0
Compare
|
Rebased on master and reduced to the keep_first one-liner: the mtmd.cpp hunk no longer built since entries are stored by value, and the vision test was dropped because no current fixture can reach a cut between two adjacent media chunks with a reused cache. This PR fixes #29866 |
* server: reject partial media truncation * server: keep only the keep_first fix Drop the mtmd test helper change, which no longer builds since clip_image_f32_batch stores its entries by value, and drop the vision test: no test fixture reaches a cut between two adjacent media chunks with a reused cache (tinygemma3 uses SWA and wraps images in text tokens, tinyopenjev and small-test are recurrent), so the test passed or failed independently of the fix. --------- Co-authored-by: Pascal <admin@serveurperso.com> (cherry picked from commit 8e16421)
Summary
server_tokens::keep_firstvalidate the cut position itself when both sides are media tokensserver_tokensusing the mtmd test chunksserver_tokenscopies media chunks internallyFixes #24075.
Edit: Fixes #29866
Tests
git diff --checkcmake -S . -B build-codex -G Ninja -DLLAMA_BUILD_TESTS=ON -DLLAMA_BUILD_SERVER=ON -DLLAMA_BUILD_EXAMPLES=OFF -DLLAMA_BUILD_TOOLS=ON -DLLAMA_BUILD_APP=OFF -DLLAMA_BUILD_UI=OFF -DLLAMA_CURL=OFF -DCMAKE_C_FLAGS="-D_WIN32_WINNT=0x0A00" -DCMAKE_CXX_FLAGS="-D_WIN32_WINNT=0x0A00"cmake --build build-codex --target test-server-tokens test-mtmd-c-api -j 2cmake -E env "PATH=C:\mingw64\bin;$env:PATH" ctest --test-dir build-codex -R "test-server-tokens|test-mtmd-c-api" --output-on-failure