iciclab

LMSM Dual-Camera Speech and Interaction Dataset

A multimodal dataset containing 50 synchronized recording sessions combining first-person and third-person video, human speech, and text labels from read-aloud scripts. The dataset captures human-robot interaction across five physical actions: press, pull, slide, twist, and insert.

Downloads146
Episodes50

Why This Matters for Physical AI

This multimodal dataset of synchronized dual-camera video, speech, and text labels supports research in embodied AI, human-robot interaction understanding, and grounding language in physical actions.

Technical Profile

Modalities
rgbaudiolanguage
Environment
lab
Task Types
human-robot-interactionpresspullslidetwistinsert
Episodes
50
Data Format
JPEG sequences, MP4 video, WAV audio, JSON, CSV
Annotation Types
language_instructionsaction_labelsprosody_labels
Part of the LMSM Dual-Camera Speech and Interaction Dataset family

Access

Need custom rgb data?

Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets