Skip to content

Video input > ~10 s hangs llama-server forever, no response and no error (deadlock in mmproj video probe) #27587

Description

@shenxchen

title: Video input > ~10 s hangs llama-server forever, no response and no error (deadlock in mmproj video probe)
labels: bug, video, http server

Environment

  • llama.cpp b10573
  • OS: Windows 11(llama-server.exe, prebuilt CUDA 13.3)
  • ffmpeg / ffprobe available in PATH

Bug summary

A video attached via input_video(OpenAI-compatible /v1/chat/completions)hangs the request forever once the video is longer than ~10–13 s(≈300–400 frames). The server itself stays healthy(other requests are answered), but this one returns nothing: no HTTP response, no error body, /slots stays idle, and the client must rely on its own timeout.

Crucially this is not a "context too large" case: the video frames are sampled at ~4 fps, so both the passing and the hanging video produce a similarly tiny prompt token count(~4750 tokens for 640×360@30fps),far below n_ctx.

Minimal reproduction

Build a server with any multimodal model(llava / Qwen-VL, requires --mmproj), e.g.:

llama-server -m model.gguf --mmproj mmproj.gguf --ctx-size 65536 --n-gpu-layers 99 --port 8080

Generate two test videos:

ffmpeg -f lavfi -i testsrc=duration=10    -c:v libx264 -pix_fmt yuv420p -movflags +faststart ok.mp4
ffmpeg -f lavfi -i testsrc=duration=13.5 -c:v libx264 -pix_fmt yuv420p -movflags +faststart hang.mp4

POST /v1/chat/completions with:

{
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "describe the video"},
      {"type": "input_video", "input_video": {"data": "<base64 of the mp4>"}}
    ]
  }],
  "max_tokens": 64
}

Result:

  • ok.mp4(10 s, ~300 frames)→ 200 OK, usage.prompt_tokens ≈ 4750
  • hang.mp4(13.5 s, ~400 frames, ~87 KB)→ request never returns

What happens under the hood

With -lv 5 the last log line printed before the hang is:

D probe: launching: ffprobe -v quiet -show_entries stream=width,height,r_frame_rate,nb_frames,duration -select_streams v:0 -of default=noprint_wrappers=1 pipe:0

After that: nothing. No progress lines, /slots stays is_processing=false, and at the moment of the hang no ffprobe/ffmpeg child process is alive in the process table — yet the request thread never wakes up.

Root-cause hypothesis(tools/mtmd/mtmd-helper.cpp)

mtmd_helper_video::probe() launches ffprobe -i pipe:0 and then runs a feeder thread (start_feeder) that fwrites the entire video buffer into the child's stdin, while the main thread blocks in while (fgets(...)) reading the child's stdout:

  • Both streams are Windows anonymous pipes with small buffers.
  • ffprobe reads the non-seekable pipe:0 and emits its final -show_entries report only after consuming enough input(no streaming output).
  • For inputs above a threshold(≈ > 400 frames / > ~10 s / > ~85 KB)the write side fills the pipe and blocks waiting for the child to read; the child does not write stdout until it has read enough; the main thread blocks on stdout ⇒ a write/read mutual-wait with no timeout.

Smaller inputs drain in time and succeed.

Why it’s the caller, not ffmpeg/ffprobe

All of the following finish instantly for both files:

  • cat ok.mp4 | ffprobe ... -i pipe:0
  • cat hang.mp4 | ffprobe ... -i pipe:0(both dump the same metadata in ~0 s)
  • cat hang.mp4 | ffmpeg -nostdin -i cache:pipe:0 -vf fps=4 -f rawvideo -pix_fmt rgb24 pipe:1
  • A Python subprocess re-implementation of the same "writer-thread + blocking reader" pattern.

So the hang is specific to the combination inside llama.cpp(sheredom anonymous pipes + CRT buffered fwrite of the whole buffer + concurrent blocking reader), not to ffmpeg/ffprobe itself.

Suggested fixes

  1. In probe() (and start_ffmpeg), write the full buffer and close stdin before reading stdout (serialize instead of concurrent blocking), or
  2. Add a timeout around the subprocess I/O and fail with a clear error instead of hanging forever, or
  3. Prefer feeding the child a temp file / file path over pipe:0 for large inputs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions