yuanzhangMIT

GEAR-VLA Dataset

Image-language dataset for object affordance-area grounding, where models identify interaction regions as bounding boxes given an image and natural-language affordance query.

Downloads0
Episodes12448
Likes1

Why This Matters for Physical AI

This dataset grounds visual affordances to natural language queries, enabling embodied AI systems to understand where and how to interact with objects in their environment.

Technical Profile

Modalities
rgblanguage
Task Types
affordance-groundingvisual-question-answering
Episodes
12448
Data Format
Parquet
Annotation Types
language_instructionsbounding_boxes
License
MIT
Part of the GEAR-VLA family

Access

Need custom rgb data?

Claru builds purpose-built datasets for any environment applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets