Sterzhang2024cc-by-nc-4.0
Video2Skill Benchmark
A benchmark for Streaming Embodied Skill Discovery (SESD) that evaluates how Vision-Language Models organize manipulation events from video streams into reusable skills, covering robot tabletop manipulation and human kitchen activity.
Downloads59
Episodes172
Why This Matters for Physical AI
This benchmark advances embodied AI by measuring how well vision-language models can discover and organize reusable manipulation skills from continuous video streams, a key capability for building generalizable robotic agents.
Technical Profile
- Modalities
- rgblanguage
- Robot Embodiments
- robot manipulatorhuman
- Environment
- kitchenlab
- Task Types
- manipulationskill_discoverytemporal_groundingevent_classification
- Episodes
- 172
- Data Format
- JSONL
- Annotation Types
- language_instructionstemporal_groundingaction_labelsskill_definitions
- License
- cc-by-nc-4.0
Access
Need custom rgb data?
Claru builds purpose-built datasets for kitchen applications with dense human annotations and quality assurance.
Request a Sample Pack