Rithvik762derived-from-navila-r2r-rxr
VLN Trajectory-Memory Stage 2 (projector alignment)
Text-only question-answering dataset where answers are derived solely from robot action histories represented as sequences of primitive navigation actions. Built to evaluate whether vision-language models can process trajectory memory injected via a trained projector.
Downloads39
Episodes24000
Why This Matters for Physical AI
This dataset evaluates how well vision-language models can maintain and reason about spatial memory from action sequences, a critical capability for embodied AI agents navigating real environments.
Technical Profile
- Modalities
- language
- Action Space
- discrete_navigation_primitives
- Environment
- simulation
- Task Types
- navigationquestion_answering
- Episodes
- 24000
- Data Format
- parquet
- Annotation Types
- language_instructionsaction_labelsquestion_answer_pairs
- License
- derived-from-navila-r2r-rxr
Access
Need custom language data?
Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.
Request a Sample Pack