xlangaiMIT

RoboFine-Bench

A fine-grained robotic video understanding benchmark for evaluating whether Vision-Language Models can capture execution-level details of robot manipulation across 500 held-out manipulation videos from 10 robot datasets. It contains 10,816 atomic facts across ten action-relevant dimensions with VQA and caption evaluation tracks.

Downloads2K
Likes5

Technical Profile

Modalities
rgblanguage
Robot Embodiments
mobile_manipulatorhumanoidquadruped
Environment
labhomekitchen
Task Types
manipulationgraspingpick_and_placepouring
Data Format
JSON
License
MIT
Part of the RoboFine-Bench family

Community Signals

Top 25% by downloads

Access

Need custom rgb data?

Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets