title: Video input > ~10 s hangs llama-server forever, no response and no error (deadlock in mmproj video probe)
labels: bug, video, http server
Environment
- llama.cpp b10573
- OS: Windows 11(llama-server.exe, prebuilt CUDA 13.3)
- ffmpeg / ffprobe available in
PATH
Bug summary
A video attached via input_video(OpenAI-compatible /v1/chat/completions)hangs the request forever once the video is longer than ~10–13 s(≈300–400 frames). The server itself stays healthy(other requests are answered), but this one returns nothing: no HTTP response, no error body, /slots stays idle, and the client must rely on its own timeout.
Crucially this is not a "context too large" case: the video frames are sampled at ~4 fps, so both the passing and the hanging video produce a similarly tiny prompt token count(~4750 tokens for 640×360@30fps),far below n_ctx.
Minimal reproduction
Build a server with any multimodal model(llava / Qwen-VL, requires --mmproj), e.g.:
llama-server -m model.gguf --mmproj mmproj.gguf --ctx-size 65536 --n-gpu-layers 99 --port 8080
Generate two test videos:
ffmpeg -f lavfi -i testsrc=duration=10 -c:v libx264 -pix_fmt yuv420p -movflags +faststart ok.mp4
ffmpeg -f lavfi -i testsrc=duration=13.5 -c:v libx264 -pix_fmt yuv420p -movflags +faststart hang.mp4
POST /v1/chat/completions with:
{
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "describe the video"},
{"type": "input_video", "input_video": {"data": "<base64 of the mp4>"}}
]
}],
"max_tokens": 64
}
Result:
ok.mp4(10 s, ~300 frames)→ 200 OK, usage.prompt_tokens ≈ 4750
hang.mp4(13.5 s, ~400 frames, ~87 KB)→ request never returns
What happens under the hood
With -lv 5 the last log line printed before the hang is:
D probe: launching: ffprobe -v quiet -show_entries stream=width,height,r_frame_rate,nb_frames,duration -select_streams v:0 -of default=noprint_wrappers=1 pipe:0
After that: nothing. No progress lines, /slots stays is_processing=false, and at the moment of the hang no ffprobe/ffmpeg child process is alive in the process table — yet the request thread never wakes up.
Root-cause hypothesis(tools/mtmd/mtmd-helper.cpp)
mtmd_helper_video::probe() launches ffprobe -i pipe:0 and then runs a feeder thread (start_feeder) that fwrites the entire video buffer into the child's stdin, while the main thread blocks in while (fgets(...)) reading the child's stdout:
- Both streams are Windows anonymous pipes with small buffers.
- ffprobe reads the non-seekable
pipe:0 and emits its final -show_entries report only after consuming enough input(no streaming output).
- For inputs above a threshold(≈ > 400 frames / > ~10 s / > ~85 KB)the write side fills the pipe and blocks waiting for the child to read; the child does not write stdout until it has read enough; the main thread blocks on stdout ⇒ a write/read mutual-wait with no timeout.
Smaller inputs drain in time and succeed.
Why it’s the caller, not ffmpeg/ffprobe
All of the following finish instantly for both files:
cat ok.mp4 | ffprobe ... -i pipe:0
cat hang.mp4 | ffprobe ... -i pipe:0(both dump the same metadata in ~0 s)
cat hang.mp4 | ffmpeg -nostdin -i cache:pipe:0 -vf fps=4 -f rawvideo -pix_fmt rgb24 pipe:1
- A Python
subprocess re-implementation of the same "writer-thread + blocking reader" pattern.
So the hang is specific to the combination inside llama.cpp(sheredom anonymous pipes + CRT buffered fwrite of the whole buffer + concurrent blocking reader), not to ffmpeg/ffprobe itself.
Suggested fixes
- In
probe() (and start_ffmpeg), write the full buffer and close stdin before reading stdout (serialize instead of concurrent blocking), or
- Add a timeout around the subprocess I/O and fail with a clear error instead of hanging forever, or
- Prefer feeding the child a temp file / file path over
pipe:0 for large inputs.
title: Video input > ~10 s hangs llama-server forever, no response and no error (deadlock in mmproj video probe)
labels: bug, video, http server
Environment
PATHBug summary
A video attached via
input_video(OpenAI-compatible/v1/chat/completions)hangs the request forever once the video is longer than ~10–13 s(≈300–400 frames). The server itself stays healthy(other requests are answered), but this one returns nothing: no HTTP response, no error body,/slotsstays idle, and the client must rely on its own timeout.Crucially this is not a "context too large" case: the video frames are sampled at ~4 fps, so both the passing and the hanging video produce a similarly tiny prompt token count(~4750 tokens for 640×360@30fps),far below
n_ctx.Minimal reproduction
Build a server with any multimodal model(llava / Qwen-VL, requires
--mmproj), e.g.:Generate two test videos:
POST /v1/chat/completionswith:{ "messages": [{ "role": "user", "content": [ {"type": "text", "text": "describe the video"}, {"type": "input_video", "input_video": {"data": "<base64 of the mp4>"}} ] }], "max_tokens": 64 }Result:
ok.mp4(10 s, ~300 frames)→200 OK,usage.prompt_tokens ≈ 4750hang.mp4(13.5 s, ~400 frames, ~87 KB)→ request never returnsWhat happens under the hood
With
-lv 5the last log line printed before the hang is:After that: nothing. No progress lines,
/slotsstaysis_processing=false, and at the moment of the hang noffprobe/ffmpegchild process is alive in the process table — yet the request thread never wakes up.Root-cause hypothesis(tools/mtmd/mtmd-helper.cpp)
mtmd_helper_video::probe()launchesffprobe -i pipe:0and then runs a feeder thread (start_feeder) thatfwrites the entire video buffer into the child's stdin, while the main thread blocks inwhile (fgets(...))reading the child's stdout:pipe:0and emits its final-show_entriesreport only after consuming enough input(no streaming output).Smaller inputs drain in time and succeed.
Why it’s the caller, not ffmpeg/ffprobe
All of the following finish instantly for both files:
cat ok.mp4 | ffprobe ... -i pipe:0cat hang.mp4 | ffprobe ... -i pipe:0(both dump the same metadata in ~0 s)cat hang.mp4 | ffmpeg -nostdin -i cache:pipe:0 -vf fps=4 -f rawvideo -pix_fmt rgb24 pipe:1subprocessre-implementation of the same "writer-thread + blocking reader" pattern.So the hang is specific to the combination inside llama.cpp(sheredom anonymous pipes + CRT buffered
fwriteof the whole buffer + concurrent blocking reader), not to ffmpeg/ffprobe itself.Suggested fixes
probe()(andstart_ffmpeg), write the full buffer and close stdin before reading stdout (serialize instead of concurrent blocking), orpipe:0for large inputs.