bhatvineet

mg-vqa-bench

Contains 600 frozen visual question answering tasks for cluttered tabletop scenes where vision-language models answer questions using single images, perception tools, or perception combined with robot manipulation.

Downloads477
Episodes600

Why This Matters for Physical AI

This dataset enables research in embodied visual reasoning by combining vision-language models with simulated robotic perception and manipulation in complex tabletop environments.

Technical Profile

Modalities
rgblanguage
Environment
simulation
Task Types
visual-question-answeringperception
Episodes
600
Data Format
json
Annotation Types
language_instructionsreward_labels
Part of the MG-VQA-Sim family

Access

Need custom rgb data?

Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets