GeWu-Lab2026
ROMA
An LLM-based system for real-world object-centric multi-sensory active perception that integrates vision, audio, touch, and force feedback through a reasoning-interaction-feedback loop. The dataset includes the ROMI-2K collection covering nearly 2,000 objects with 6 atomic interactions and synchronized multi-sensory feedback.
Downloads6
Episodes~2000 objects
Likes1
Why This Matters for Physical AI
ROMA demonstrates how multi-sensory perception and active reasoning through language models can enable robots to understand object properties through diverse interactions, advancing embodied AI systems that can reason about and explore their environment.
Technical Profile
- Modalities
- rgbaudiotactileforce_torque
- Robot Embodiments
- robotic_armhandheld
- Action Space
- interaction_selection
- Environment
- labtabletop
- Task Types
- object_interactionactive_perceptionmanipulation
- Episodes
- ~2000 objects
- Annotation Types
- action_labelsbounding_boxesmodality_labels
Access
Need custom rgb data?
Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.
Request a Sample Pack