RLE-Bench

Codex Benchmark

A benchmark dataset of 42 RoboDojo tasks with 31 successful completions, evaluated using Codex CLI with GPT-6 Astra in stepped simulation environments with native spectator recordings.

Downloads63
Episodes42

Why This Matters for Physical AI

This benchmark evaluates large language model-based robot control systems across diverse simulated manipulation tasks, providing standardized metrics for assessing embodied AI performance.

Technical Profile

Modalities
rgb
Environment
simulation
Episodes
42
Annotation Types
language_instructionsaction_labels
Part of the RoboDojo family

Access

Need custom rgb data?

Claru builds purpose-built datasets for simulation applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets