irl-kit2026

SPARC VQA Raw

Unfiltered corpus of 838,211 vision-language examples with embedded images, questions, answers, and spatial reasoning annotations from robot demonstrations. Includes a filtered subset of 284,909 examples used for training Qwen3.5 models.

Downloads0
Episodes838211

Why This Matters for Physical AI

SPARC VQA provides large-scale spatial reasoning and visual understanding annotations grounded in robot demonstrations, enabling vision-language models to learn embodied spatial concepts essential for robotic manipulation and navigation tasks.

Technical Profile

Modalities
rgblanguage
Task Types
spatial_reasoningvisual_question_answering
Episodes
838211
Data Format
Parquet
Annotation Types
language_instructionslanguage_answersspatial_annotationsquality_scores
Part of the SPARC family

Access

Need custom rgb data?

Claru builds purpose-built datasets for any environment applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets