xlangaiMIT
RoboFine-Bench
A fine-grained robotic video understanding benchmark for evaluating whether Vision-Language Models can capture execution-level details of robot manipulation across 500 held-out manipulation videos from 10 robot datasets. It contains 10,816 atomic facts across ten action-relevant dimensions with VQA and caption evaluation tracks.
Downloads2K
Likes5
Technical Profile
- Modalities
- rgblanguage
- Robot Embodiments
- mobile_manipulatorhumanoidquadruped
- Environment
- labhomekitchen
- Task Types
- manipulationgraspingpick_and_placepouring
- Data Format
- JSON
- License
- MIT
Community Signals
Top 25% by downloads
HuggingFace Discussions1
Access
Need custom rgb data?
Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.
Request a Sample Pack