RLE-Bench

kinex (v0.6.0) vs kinex (v0.6.0-improve) vs codex · LIBERO Long · GPT-6 Astra / medium

Evaluation dataset comparing three agent variants (kinex v0.6.0, kinex v0.6.0-improve, and codex) on LIBERO Long tasks using GPT-6 Astra medium model with 45 planned episodes.

Downloads673
Episodes45

Why This Matters for Physical AI

This evaluation dataset enables comparative analysis of different embodied AI agent implementations on standardized manipulation benchmarks, supporting research into agent design improvements and model capabilities for robotic task execution.

Technical Profile

Episodes
45
Annotation Types
action_labelsreward_labels
Part of the LIBERO family

Access

Need custom physical AI data?

Claru builds purpose-built datasets for any environment applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets