GeWu-Lab2026

ROMA

An LLM-based system for real-world object-centric multi-sensory active perception that integrates vision, audio, touch, and force feedback through a reasoning-interaction-feedback loop. The dataset includes the ROMI-2K collection covering nearly 2,000 objects with 6 atomic interactions and synchronized multi-sensory feedback.

Downloads6
Episodes~2000 objects
Likes1

Why This Matters for Physical AI

ROMA demonstrates how multi-sensory perception and active reasoning through language models can enable robots to understand object properties through diverse interactions, advancing embodied AI systems that can reason about and explore their environment.

Technical Profile

Modalities
rgbaudiotactileforce_torque
Robot Embodiments
robotic_armhandheld
Action Space
interaction_selection
Environment
labtabletop
Task Types
object_interactionactive_perceptionmanipulation
Episodes
~2000 objects
Annotation Types
action_labelsbounding_boxesmodality_labels
Part of the ROMA family

Access

Need custom rgb data?

Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets