Sterzhang2024cc-by-nc-4.0

Video2Skill Benchmark

A benchmark for Streaming Embodied Skill Discovery (SESD) that evaluates how Vision-Language Models organize manipulation events from video streams into reusable skills, covering robot tabletop manipulation and human kitchen activity.

Downloads59
Episodes172

Why This Matters for Physical AI

This benchmark advances embodied AI by measuring how well vision-language models can discover and organize reusable manipulation skills from continuous video streams, a key capability for building generalizable robotic agents.

Technical Profile

Modalities
rgblanguage
Robot Embodiments
robot manipulatorhuman
Environment
kitchenlab
Task Types
manipulationskill_discoverytemporal_groundingevent_classification
Episodes
172
Data Format
JSONL
Annotation Types
language_instructionstemporal_groundingaction_labelsskill_definitions
License
cc-by-nc-4.0
Part of the Video2Skill family

Access

Need custom rgb data?

Claru builds purpose-built datasets for kitchen applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets