Myungkyu2026

RoboDojo-taco-visual2-gemini

A visual-grounding variant of RoboDojo-taco-gemini featuring 800 long-horizon episodes across 8 tasks with dense high-level labels and point markers drawn into head camera video frames to indicate target positions for low-level policy learning.

Downloads80
Episodes800

Why This Matters for Physical AI

This dataset enables training of visual grounding policies that learn to interpret spatial instructions by following point markers in egocentric camera views, advancing embodied AI systems' ability to understand and execute spatially-grounded natural language commands.

Technical Profile

Modalities
rgbproprioception
Robot Embodiments
dual ARX X5
Action Space
joint_positions
Environment
simulation
Task Types
manipulationpick_and_placeobject_classificationstackingtic_tac_toe
Episodes
800
Data Format
LeRobot
Annotation Types
language_instructionsaction_labelsvisual_groundingbounding_boxes
Part of the RoboDojo family

Access

Need custom rgb data?

Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets