yuanzhangMIT
GEAR-VLA Dataset
Image-language dataset for object affordance-area grounding, where models identify interaction regions as bounding boxes given an image and natural-language affordance query.
Downloads0
Episodes12448
Likes1
Why This Matters for Physical AI
This dataset grounds visual affordances to natural language queries, enabling embodied AI systems to understand where and how to interact with objects in their environment.
Technical Profile
- Modalities
- rgblanguage
- Task Types
- affordance-groundingvisual-question-answering
- Episodes
- 12448
- Data Format
- Parquet
- Annotation Types
- language_instructionsbounding_boxes
- License
- MIT
Access
Need custom rgb data?
Claru builds purpose-built datasets for any environment applications with dense human annotations and quality assurance.
Request a Sample Pack