feat(perf): add real multi-image MMMU dataset mode - #1729
Conversation
| image = self._to_pil_image(item.get(f'image_{index}')) | ||
| if image is None: | ||
| continue | ||
| image_urls.append(PIL_to_base64(image, add_header=True)) |
There was a problem hiding this comment.
This fails against the documented default dataset configuration. PIL_to_base64(..., add_header=True) defaults to JPEG and saves the image without a mode conversion, but the real AI-ModelScope/MMMU Music validation rows contain RGBA images (including all five images in the first row). Pillow raises OSError: cannot write mode RGBA as JPEG, so build_messages() cannot produce even its first request.
Please normalize the image before encoding (for example, explicitly flatten/convert alpha-bearing modes such as RGBA/LA/P to RGB before the existing JPEG path), or encode those images as PNG and ensure the data-URL MIME type matches. Please also add a regression test using an RGBA image and exercise build_messages() end-to-end; the current RGB-only fixture cannot detect this failure.
Yunnglin
left a comment
There was a problem hiding this comment.
Request changes: please address the blocking inline review finding before merge.
|
按照line_by_line格式读取的请求中,image_url是否支持本地路径? |
|
当前 |
Yunnglin
left a comment
There was a problem hiding this comment.
LGTM — multi-image workload now round-robins all MMMU subjects; targeted coverage and all CI checks pass.
Closes #1726.
What
Adds a discoverable, real-data multi-image mode for
evalscope perf: a newmmmu_multi_imagedataset plugin that builds multi-image stress requests from the open-source MMMU validation set, plus EN/ZH user docs for multi-image stress testing.Why this instead of existing paths
create_message) already supports multipleimage_urlparts in one message, andline_by_linealready accepts OpenAI-style messages arrays / full request bodies, so private or constructed multi-image payloads keep working there unchanged.How it works
evalscope/perf/plugin/datasets/mmmu_multi_image.py(registered asmmmu_multi_image): loads MMMUvalidationfor the configuredsubset(defaultMusic), collects non-emptyimage_1…image_7fields in source order (PIL /bytes-dict /path-dict representations), keeps rows with at leastmin_imagesimages, and yields one user message per row via the sharedcreate_message(text, image_urls=[...])path.MMMUMultiImageDatasetArgs:subset: str = 'Music',min_images: int = 2validated to 2–7 (MMMU rows carry at most 7 images).--tokenize-promptis rejected with a clear error: multimodal messages cannot be represented as a flat token-ID list.docs/{en,zh}/user_guides/stress_test/multi_image.md(MMMU mode +line_by_linecustom-data recipe), linked from both stress-test indexes.evalscope eval --datasets mmmufor official scoring.evalscope perf \ --model your-vl-model \ --url http://localhost:8000/v1/chat/completions \ --dataset mmmu_multi_image \ --dataset-args '{"subset":"Music","min_images":2}' \ --parallel 4 \ --number 100Validation performed
pytest tests/perf/test_mmmu_multi_image.py -q→ 4 passed (message structure/order, min-image filtering, arg validation, tokenize-mode rejection; hub loading mocked, no network).ruff checkandruff format --check(repo-pinned ruff 0.16.4) on the touched Python files → clean.pytest tests/perf --collect-only→ 370 collected, no registration breakage.git diff --check→ clean.