RLE-Bench

Kinex (v0.10.0) Benchmark

A robotics evaluation benchmark consisting of 82 selected tasks across 5 environments with 72 recorded results and 10 pending tasks. Includes original and modified instructions with metrics for outcome, execution, evidence validity, and usage coverage.

Downloads646
Episodes82

Why This Matters for Physical AI

Provides a structured evaluation benchmark for assessing robotics task performance across diverse environments and instruction variants, enabling standardized comparison of embodied AI systems.

Technical Profile

Episodes
82
Annotation Types
language_instructionsreward_labels
Part of the RLE-Bench family

Access

Need custom physical AI data?

Claru builds purpose-built datasets for any environment applications with dense human annotations and quality assurance.

Request a Sample Pack