qnguyen3MIT

Crosswire

A video question-answering benchmark testing whether multimodal models can identify which robot performed which action when multiple Franka robots share a frame. Contains 650 multiple-choice questions over 497 videos of LIBERO manipulation tasks in three difficulty tiers.

Downloads0
Episodes497

Why This Matters for Physical AI

This benchmark tests multimodal models' ability to understand and distinguish between multiple robot agents performing manipulation tasks, which is critical for embodied AI systems that must operate in multi-agent environments and understand scene dynamics from video.

Technical Profile

Modalities
rgblanguage
Robot Embodiments
Franka Panda
Environment
simulation
Task Types
manipulationvideo-question-answering
Episodes
497
Data Format
JSONL
Annotation Types
language_instructionsmultiple_choice_qatask_labels
License
MIT
Part of the LIBERO family

Access

Need custom rgb data?

Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets