RLE-Bench
Kinex (v0.10.0) Benchmark
A robotics evaluation benchmark consisting of 82 selected tasks across 5 environments with 72 recorded results and 10 pending tasks. Includes original and modified instructions with metrics for outcome, execution, evidence validity, and usage coverage.
Downloads646
Episodes82
Why This Matters for Physical AI
Provides a structured evaluation benchmark for assessing robotics task performance across diverse environments and instruction variants, enabling standardized comparison of embodied AI systems.
Technical Profile
- Episodes
- 82
- Annotation Types
- language_instructionsreward_labels
Access
Need custom physical AI data?
Claru builds purpose-built datasets for any environment applications with dense human annotations and quality assurance.
Request a Sample Pack