deepannotateaicc-by-4.0
Vegetable Harvesting Multimodal Dataset v1
A first-person egocentric dataset capturing hands, objects, poses, and activity annotations for agricultural manual tasks including vegetable harvesting and plucking. Contains 200 annotated frames with multimodal data including RGB video, depth, hand keypoints, object bounding boxes, and natural language action descriptions.
Downloads1
Episodes200
Hours0.38
Why This Matters for Physical AI
This dataset provides egocentric hand-object interaction data for agricultural robotics and embodied AI research, enabling systems to learn fine-grained manipulation skills and human labor mechanics from multimodal observations including hand pose, object detection, and semantic action descriptions.
Technical Profile
- Modalities
- rgbdepthhand_keypoints_2dhand_keypoints_3dimulanguagetrajectoriesbounding_boxes
- Action Space
- language
- Environment
- agriculture
- Task Types
- manipulationgraspingpick_and_placehuman-object-interaction
- Episodes
- 200
- Total Hours
- 0.38
- Data Format
- HDF5
- Annotation Types
- hand_keypoints_2dhand_keypoints_3dhand_object_interactionsbounding_boxesaction_labelslanguage_instructionsmotion_statisticstrajectories
- License
- cc-by-4.0
Access
Need custom rgb data?
Claru builds purpose-built datasets for agriculture applications with dense human annotations and quality assurance.
Request a Sample Pack