AIBenchcc-by-4.0

M-VQA: Visual Question Answering under Image Distortions

A dataset of 12,400 image-question-answer samples with 400 original images and 12,000 distorted variants across 30 distortion types and 5 severity levels, designed to evaluate the robustness of multimodal large language models in visual question answering under image degradation.

Downloads81

Technical Profile

Modalities
rgblanguage
Task Types
visual-question-answering
Data Format
TSV
License
cc-by-4.0
Part of the M-VQA: Visual Question Answering under Image Distortions family

Access

Need custom rgb data?

Claru builds purpose-built datasets for any environment applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets