iciclab
LMSM Dual-Camera Speech and Interaction Dataset
A multimodal dataset containing 50 synchronized recording sessions combining first-person and third-person video, human speech, and text labels from read-aloud scripts. The dataset captures human-robot interaction across five physical actions: press, pull, slide, twist, and insert.
Downloads146
Episodes50
Why This Matters for Physical AI
This multimodal dataset of synchronized dual-camera video, speech, and text labels supports research in embodied AI, human-robot interaction understanding, and grounding language in physical actions.
Technical Profile
- Modalities
- rgbaudiolanguage
- Environment
- lab
- Task Types
- human-robot-interactionpresspullslidetwistinsert
- Episodes
- 50
- Data Format
- JPEG sequences, MP4 video, WAV audio, JSON, CSV
- Annotation Types
- language_instructionsaction_labelsprosody_labels
Access
Need custom rgb data?
Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.
Request a Sample Pack