naishashettymit

MILO Benchmark v1.0

A small, versioned dataset of (scene, instruction, ground-truth task spec) triples for evaluating embodied task planning in AI2-THOR, comprising 25 tasks across 5 scenes with natural-language instructions paired with structured goal specifications.

Downloads0
Episodes25

Why This Matters for Physical AI

This dataset provides a structured benchmark for evaluating vision-language-robotics task planning in embodied AI agents, with explicit success predicates and known limitations documented for robust evaluation.

Technical Profile

Modalities
language
Robot Embodiments
embodied_agent
Environment
simulation
Task Types
object_localizationnavigationgraspingpick_and_placecontainer_interaction
Episodes
25
Data Format
json
Annotation Types
language_instructionstask_specificationssuccess_predicates
License
mit
Part of the MILO family

Access

Need custom language data?

Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets