RLE-Bench
Codex Benchmark
A benchmark dataset of 42 RoboDojo tasks with 31 successful completions, evaluated using Codex CLI with GPT-6 Astra in stepped simulation environments with native spectator recordings.
Downloads63
Episodes42
Why This Matters for Physical AI
This benchmark evaluates large language model-based robot control systems across diverse simulated manipulation tasks, providing standardized metrics for assessing embodied AI performance.
Technical Profile
- Modalities
- rgb
- Environment
- simulation
- Episodes
- 42
- Annotation Types
- language_instructionsaction_labels
Access
Need custom rgb data?
Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.
Request a Sample Pack