Skip to content

feat(perf): add real multi-image MMMU dataset mode - #1729

Merged
Yunnglin merged 23 commits into
modelscope:mainfrom
Bruce-Yii:feat/perf-multi-image-mmmu-1726
Sep 17, 2026
Merged

Yunnglin merged 23 commits into
modelscope:mainfrom
Bruce-Yii:feat/perf-multi-image-mmmu-1726

Conversation

@Bruce-Yii

Copy link
Copy Markdown
Contributor

Closes #1726.

What

Adds a discoverable, real-data multi-image mode for evalscope perf: a new mmmu_multi_image dataset plugin that builds multi-image stress requests from the open-source MMMU validation set, plus EN/ZH user docs for multi-image stress testing.

Why this instead of existing paths

  • The perf transport (create_message) already supports multiple image_url parts in one message, and line_by_line already accepts OpenAI-style messages arrays / full request bodies, so private or constructed multi-image payloads keep working there unchanged.
  • What was missing (issue 希望在压测的多模态数据集类型中加入支持一次输入多张图片的数据集模式 #1726) is a built-in, non-random, open-source multi-image dataset mode plus documentation of the custom-data path. This PR fills exactly that gap without duplicating the generic transport path.

How it works

  • evalscope/perf/plugin/datasets/mmmu_multi_image.py (registered as mmmu_multi_image): loads MMMU validation for the configured subset (default Music), collects non-empty image_1 … image_7 fields in source order (PIL / bytes-dict / path-dict representations), keeps rows with at least min_images images, and yields one user message per row via the shared create_message(text, image_urls=[...]) path.
  • MMMUMultiImageDatasetArgs: subset: str = 'Music', min_images: int = 2 validated to 2–7 (MMMU rows carry at most 7 images).
  • --tokenize-prompt is rejected with a clear error: multimodal messages cannot be represented as a flat token-ID list.
  • Docs: new docs/{en,zh}/user_guides/stress_test/multi_image.md (MMMU mode + line_by_line custom-data recipe), linked from both stress-test indexes.
  • This is stress/performance traffic support, not MMMU benchmark scoring — use evalscope eval --datasets mmmu for official scoring.
evalscope perf \
  --model your-vl-model \
  --url http://localhost:8000/v1/chat/completions \
  --dataset mmmu_multi_image \
  --dataset-args '{"subset":"Music","min_images":2}' \
  --parallel 4 \
  --number 100

Validation performed

  • pytest tests/perf/test_mmmu_multi_image.py -q → 4 passed (message structure/order, min-image filtering, arg validation, tokenize-mode rejection; hub loading mocked, no network).
  • ruff check and ruff format --check (repo-pinned ruff 0.16.4) on the touched Python files → clean.
  • pytest tests/perf --collect-only → 370 collected, no registration breakage.
  • git diff --check → clean.
  • Full test-suite validation was not run; no such claim is made.

image = self._to_pil_image(item.get(f'image_{index}'))
if image is None:
continue
image_urls.append(PIL_to_base64(image, add_header=True))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This fails against the documented default dataset configuration. PIL_to_base64(..., add_header=True) defaults to JPEG and saves the image without a mode conversion, but the real AI-ModelScope/MMMU Music validation rows contain RGBA images (including all five images in the first row). Pillow raises OSError: cannot write mode RGBA as JPEG, so build_messages() cannot produce even its first request.

Please normalize the image before encoding (for example, explicitly flatten/convert alpha-bearing modes such as RGBA/LA/P to RGB before the existing JPEG path), or encode those images as PNG and ensure the data-URL MIME type matches. Please also add a regression test using an RGBA image and exercise build_messages() end-to-end; the current RGB-only fixture cannot detect this failure.

@Yunnglin Yunnglin left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Request changes: please address the blocking inline review finding before merge.

@JamiePW

JamiePW commented Sep 17, 2026

Copy link
Copy Markdown

按照line_by_line格式读取的请求中,image_url是否支持本地路径?

@Yunnglin

Copy link
Copy Markdown
Collaborator

当前 line_by_line 会原样转发 JSON。image_url.url 中的本地路径不会被 EvalScope 读取或转换;只有目标服务可访问同一文件系统且支持该路径时才可能可用。建议使用 HTTP(S) URL 或服务端支持的 data URL。

@Yunnglin Yunnglin left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — multi-image workload now round-robins all MMMU subjects; targeted coverage and all CI checks pass.

@Yunnglin
Yunnglin merged commit deac60e into modelscope:main Sep 17, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

希望在压测的多模态数据集类型中加入支持一次输入多张图片的数据集模式

3 participants