RLE-Bench
kinex (v0.6.0) vs kinex (v0.6.0-improve) vs codex · LIBERO Long · GPT-6 Astra / medium
Evaluation dataset comparing three agent variants (kinex v0.6.0, kinex v0.6.0-improve, and codex) on LIBERO Long tasks using GPT-6 Astra medium model with 45 planned episodes.
Downloads673
Episodes45
Why This Matters for Physical AI
This evaluation dataset enables comparative analysis of different embodied AI agent implementations on standardized manipulation benchmarks, supporting research into agent design improvements and model capabilities for robotic task execution.
Technical Profile
- Episodes
- 45
- Annotation Types
- action_labelsreward_labels
Access
Need custom physical AI data?
Claru builds purpose-built datasets for any environment applications with dense human annotations and quality assurance.
Request a Sample Pack