Myungkyu2026
RoboDojo-taco-visual2-gemini
A visual-grounding variant of RoboDojo-taco-gemini featuring 800 long-horizon episodes across 8 tasks with dense high-level labels and point markers drawn into head camera video frames to indicate target positions for low-level policy learning.
Downloads80
Episodes800
Why This Matters for Physical AI
This dataset enables training of visual grounding policies that learn to interpret spatial instructions by following point markers in egocentric camera views, advancing embodied AI systems' ability to understand and execute spatially-grounded natural language commands.
Technical Profile
- Modalities
- rgbproprioception
- Robot Embodiments
- dual ARX X5
- Action Space
- joint_positions
- Environment
- simulation
- Task Types
- manipulationpick_and_placeobject_classificationstackingtic_tac_toe
- Episodes
- 800
- Data Format
- LeRobot
- Annotation Types
- language_instructionsaction_labelsvisual_groundingbounding_boxes
Access
Need custom rgb data?
Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.
Request a Sample Pack