qnguyen3MIT
Crosswire
A video question-answering benchmark testing whether multimodal models can identify which robot performed which action when multiple Franka robots share a frame. Contains 650 multiple-choice questions over 497 videos of LIBERO manipulation tasks in three difficulty tiers.
Downloads0
Episodes497
Why This Matters for Physical AI
This benchmark tests multimodal models' ability to understand and distinguish between multiple robot agents performing manipulation tasks, which is critical for embodied AI systems that must operate in multi-agent environments and understand scene dynamics from video.
Technical Profile
- Modalities
- rgblanguage
- Robot Embodiments
- Franka Panda
- Environment
- simulation
- Task Types
- manipulationvideo-question-answering
- Episodes
- 497
- Data Format
- JSONL
- Annotation Types
- language_instructionsmultiple_choice_qatask_labels
- License
- MIT
Access
Need custom rgb data?
Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.
Request a Sample Pack