cq8382025cc-by-nc-4.0

LongVILBench

A benchmark for long-horizon visual imitation learning from real-world tabletop demonstration videos, containing 150 manipulation tasks with 300 demonstration videos recorded under clean and complex visual conditions.

Downloads52
Episodes150

Why This Matters for Physical AI

LongVILBench enables evaluation of vision-language models for long-horizon robotic manipulation planning by requiring inference of temporally ordered, spatially grounded action plans from human demonstrations.

Technical Profile

Modalities
rgblanguage
Action Space
language
Environment
lab
Task Types
manipulationpick_and_placegrasping
Episodes
150
Data Format
JSON
Annotation Types
language_instructionsaction_labelssemantic_action_sequences
License
cc-by-nc-4.0
Part of the LongVILBench family

Access

Need custom rgb data?

Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets