bhatvineet
mg-vqa-bench
Contains 600 frozen visual question answering tasks for cluttered tabletop scenes where vision-language models answer questions using single images, perception tools, or perception combined with robot manipulation.
Downloads477
Episodes600
Why This Matters for Physical AI
This dataset enables research in embodied visual reasoning by combining vision-language models with simulated robotic perception and manipulation in complex tabletop environments.
Technical Profile
- Modalities
- rgblanguage
- Environment
- simulation
- Task Types
- visual-question-answeringperception
- Episodes
- 600
- Data Format
- json
- Annotation Types
- language_instructionsreward_labels
Access
Need custom rgb data?
Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.
Request a Sample Pack