Skip to content

llama-context : sync pending async copies before clearing embd_seq - #25676

Merged
ggerganov merged 1 commit into
ggml-org:masterfrom
o7si:issue-25654
Jul 30, 2026
Merged

ggerganov merged 1 commit into
ggml-org:masterfrom
o7si:issue-25654

Conversation

@o7si

@o7si o7si commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Fixes #25654

Pooled embeddings (--pooling last/mean/cls) are written into the per-sequence embd_seq buffers with an asynchronous backend copy. embd_seq is freed at the start of the next decode() / encode() (both clear it) and by the destructor, without waiting for the copy.

Without that synchronization, the copy can complete after the buffer is freed, writing into freed memory and corrupting the allocator's metadata, which aborts the process on the next allocation:

llama-server(89298,0x1f9140c00) malloc: Incorrect checksum for freed object 0x13c888c00: probably modified after being freed.
Corrupt value: 0xc09339813e8a7d6f

This PR synchronizes before embd_seq is cleared in decode() / encode(), and in the destructor before the output buffers are freed.

Requirements

@CoruNethron

Copy link
Copy Markdown

Tested according to #25654 , the issue is not reproducible anymore in my setup.
Thank's for your effort, @o7si ; Fast and precise. 👍

@o7si

o7si commented Jul 15, 2026

Copy link
Copy Markdown
Contributor Author

Tested on macOS (Apple M5, Metal) with Qwen3-VL-Embedding-2B (Q6_K + f16 mmproj). Following the repro from the issue (#25654), I hammered the endpoint with sequential and concurrent image embedding requests, no abort:

cat example.jpg | base64 > ex.b64

REQ='{"input":[{"prompt_string":"<__media__>","multimodal_data":["'$(cat ex.b64)'"]}]}'

# 60 sequential requests
for i in $(seq 1 60); do
  curl -s -X POST "http://127.0.0.1:8080/v1/embeddings" \
    -H "Content-Type: application/json" -d "$REQ" \
    -o /dev/null -w "seq $i: %{http_code}\n"
done

# 8 concurrent requests
for i in $(seq 1 8); do
  curl -s -X POST "http://127.0.0.1:8080/v1/embeddings" \
    -H "Content-Type: application/json" -d "$REQ" \
    -o /dev/null -w "conc $i: %{http_code}\n" &
done
wait

All requests return 200 with valid embeddings and no abort.

@CoruNethron

CoruNethron commented Jul 15, 2026 •

Copy link
Copy Markdown

@o7si ,FYI, for this particular usecase it might be helpful to iterate different images for testing, cause otherwise server somehow returns cached response, without re-evaluating inference.
I tested with different images sequence and see no issues with your patch anyway.

@o7si
o7si marked this pull request as ready for review July 19, 2026 07:57
@o7si
o7si requested a review from ggerganov as a code owner July 19, 2026 07:57
@ggerganov ggerganov added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Jul 30, 2026
@ggerganov
ggerganov merged commit 432d7ff into ggml-org:master Jul 30, 2026
20 of 25 checks passed
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Aug 2, 2026
ishikawa added a commit to ishikawa/llama.cpp that referenced this pull request Aug 9, 2026
#7 (1e08ebf) で追加した encode()/decode() の embd_seq 解放ガード
(非同期出力コピー完了を待ってから embd_seq.clear() する) は、upstream
432d7ff (ggml-org#25676, "llama-context : sync pending async copies before
clearing embd_seq") で同等のガードが独立に追加され、destructor 側の
同期も含めて fork 版より広い範囲を保護している。

コード上の guard 自体 (if (!embd_seq.empty()) { synchronize(); }) は
upstream マージ後も機能的に同一のまま残っているため、revert は差分を
生まない。本コミットは fork 独自のコメント文言を upstream 432d7ff の
文言に揃えることで、fork 独自実装としての痕跡を除き、upstream 追従を
妨げないようにする。
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval bug: MTMD embedding lead to: malloc: Incorrect checksum for freed object

3 participants