mulligan2026MIT

Thread Nut: round 5 evaluation

Real-robot held-out evaluation dataset for the Mulligan paper containing episodes from evaluation sessions with multiple policy comparisons (HG-DAgger, HG-DAgger+Mulligan, HiL-IDQL+Mulligan) on thread nut tasks.

Downloads991
Episodes50

Why This Matters for Physical AI

This dataset provides real-world evaluation benchmarks for imitation learning and learning-from-human-feedback methods on robotic manipulation tasks, enabling assessment of policy performance and comparison across different training approaches.

Technical Profile

Modalities
rgb
Environment
lab
Task Types
manipulation
Episodes
50
Data Format
parquet
Annotation Types
language_instructionsreward_labelsaction_labels
License
MIT
Part of the Mulligan family

Access

Need custom rgb data?

Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets