deepannotateaicc-by-4.0

Vegetable Harvesting Multimodal Dataset v1

A first-person egocentric dataset capturing hands, objects, poses, and activity annotations for agricultural manual tasks including vegetable harvesting and plucking. Contains 200 annotated frames with multimodal data including RGB video, depth, hand keypoints, object bounding boxes, and natural language action descriptions.

Downloads1
Episodes200
Hours0.38

Why This Matters for Physical AI

This dataset provides egocentric hand-object interaction data for agricultural robotics and embodied AI research, enabling systems to learn fine-grained manipulation skills and human labor mechanics from multimodal observations including hand pose, object detection, and semantic action descriptions.

Technical Profile

Modalities
rgbdepthhand_keypoints_2dhand_keypoints_3dimulanguagetrajectoriesbounding_boxes
Action Space
language
Environment
agriculture
Task Types
manipulationgraspingpick_and_placehuman-object-interaction
Episodes
200
Total Hours
0.38
Data Format
HDF5
Annotation Types
hand_keypoints_2dhand_keypoints_3dhand_object_interactionsbounding_boxesaction_labelslanguage_instructionsmotion_statisticstrajectories
License
cc-by-4.0
Part of the Vegetable Harvesting Multimodal Dataset v1 family

Access

Need custom rgb data?

Claru builds purpose-built datasets for agriculture applications with dense human annotations and quality assurance.

Request a Sample Pack

Related Datasets