lvesucces2027apache-2.0
Safety Inspector ICRA 2027 Held-Out Test Benchmark
A held-out test benchmark containing 50 tabletop scenes and 100 query tasks for evaluating vision-language model safety inspection, rule selection, and test-time adaptation in robotic manipulation.
Downloads0
Episodes100
Why This Matters for Physical AI
This benchmark dataset evaluates vision-language models' ability to inspect robot safety and correct manipulation plans, advancing research in safe autonomous robotic systems that can reason about task constraints and adapt online.
Technical Profile
- Modalities
- rgblanguage
- Environment
- labtabletop
- Task Types
- visual-question-answeringsafety-inspectionmanipulation
- Episodes
- 100
- Annotation Types
- language_instructionsaction_labelsreward_labels
- License
- apache-2.0
Access
Need custom rgb data?
Claru builds purpose-built datasets for lab applications with dense human annotations and quality assurance.
Request a Sample Pack