Rithvik762derived-from-navila-r2r-rxr

VLN Trajectory-Memory Stage 2 (projector alignment)

Text-only question-answering dataset where answers are derived solely from robot action histories represented as sequences of primitive navigation actions. Built to evaluate whether vision-language models can process trajectory memory injected via a trained projector.

Downloads39
Episodes24000

Why This Matters for Physical AI

This dataset evaluates how well vision-language models can maintain and reason about spatial memory from action sequences, a critical capability for embodied AI agents navigating real environments.

Technical Profile

Modalities
language
Action Space
discrete_navigation_primitives
Environment
simulation
Task Types
navigationquestion_answering
Episodes
24000
Data Format
parquet
Annotation Types
language_instructionsaction_labelsquestion_answer_pairs
License
derived-from-navila-r2r-rxr
Part of the VLN Trajectory-Memory family

Access

Need custom language data?

Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets