cq8382025cc-by-nc-4.0
LongVILBench
A benchmark for long-horizon visual imitation learning from real-world tabletop demonstration videos, containing 150 manipulation tasks with 300 demonstration videos recorded under clean and complex visual conditions.
Downloads52
Episodes150
Why This Matters for Physical AI
LongVILBench enables evaluation of vision-language models for long-horizon robotic manipulation planning by requiring inference of temporally ordered, spatially grounded action plans from human demonstrations.
Technical Profile
- Modalities
- rgblanguage
- Action Space
- language
- Environment
- lab
- Task Types
- manipulationpick_and_placegrasping
- Episodes
- 150
- Data Format
- JSON
- Annotation Types
- language_instructionsaction_labelssemantic_action_sequences
- License
- cc-by-nc-4.0
Access
Need custom rgb data?
Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.
Request a Sample Pack