truelabelRequest dataEarnRequest

Physical AI Glossary

Hand-Object Interaction

Hand-object interaction (HOI) research studies how human hands contact, grasp, manipulate, and release objects across reach, grip, in-hand adjustment, and release phases. HOI datasets provide demonstration data that teaches dexterous robots to replicate human manipulation skills in unstructured environments. Leading benchmarks include EPIC-KITCHENS (100 hours of egocentric kitchen tasks), DexYCB (582,000 RGB-D frames with 3D hand pose and object pose), and DROID (76,000 trajectories across 564 scenes). Procurement requires verifying 3D hand pose accuracy, contact annotation density, object diversity, and licensing terms for commercial model training.

Updated 2026-07-1411 min read
By Truelabel Team
Reviewed by Truelabel Team ·
hand-object interaction

Quick facts

Topic
Hand Object Interaction
Audience
Procurement leads, ML ops, robotics engineers
Deliverable
Buyer-facing reference + procurement guidance

What Hand-Object Interaction Data Captures

A complete hand-object interaction (HOI) record ties four streams to one timeline: RGB-D video, 3D hand pose (a 21-joint skeleton or a MANO mesh, whose 45 pose parameters plus 10 shape and 6 for global rotation and translation rebuild a full hand), 6-DOF object pose, and a contact map naming which hand vertices touch which object surface. Task metadata (object identity, action label, success flag) turns raw motion into a supervised signal. Drop any one stream and you lose a capability: no object pose means no grasp reconstruction, no contact map means no force-closure reasoning.

No single capture method delivers all four cleanly. Lab mocap buys sub-millimeter pose at the cost of task diversity. Egocentric video buys natural context and affordances but loses 3D hand pose. Teleoperation buys diverse real scenes and end-effector trajectories but usually drops hand pose when operators drive with a joystick or VR controller instead of a mocap glove. The table below is the trade-off you are actually choosing between.

DexYCB anchors the lab-mocap corner with 582,000 RGB-D frames of 10 subjects handling 20 YCB objects, hand pose from magnetic tracking and object pose from AprilTags[1]. EPIC-KITCHENS-100 anchors egocentric with 100 hours across 700 action segments in 45 kitchens, labeled with 97 verbs and 300 nouns[2]. DROID anchors teleoperation with 76,000 trajectories (350 hours) over 564 real scenes[3].

Exocentric multi-camera rigs like H2O and First-Person Hand Action triangulate 3D pose from calibrated views but stay lab-bound. Assembly101 records synchronized egocentric and exocentric streams over 4,000 toy-assembly procedures, and Ego4D scales egocentric capture to 3,670 hours from 74 locations with 5.6 million labeled interactions[4].

Modality3D hand poseObject poseContact mapTask contextExamples
Lab mocapSub-mmYesYes (thermal/pressure)Low diversityDexYCB, GRAB
Egocentric videoNoRareNoHigh, naturalEPIC-KITCHENS, Ego4D
TeleoperationEnd-effector onlyVia robot stateGripper proxyDiverse scenesDROID, BridgeData V2
Exocentric multi-camTriangulatedYesOptionalLab-boundDexYCB, HO-3D
What each capture modality actually delivers

Detection, Recognition, and Reconstruction Tasks

HOI splits into three vision tasks that need different labels, and buyers routinely pay for the wrong one. Detection localizes interacting hand-object pairs with bounding boxes; HICO-DET (47,776 images, 150,000 pairs, 600 interaction categories) is still the standard detection benchmark[5]. Recognition classifies the interaction itself: grasp, pour, cut, twist. Reconstruction is the expensive one, recovering full 3D hand and object geometry plus the contact surface, so it needs vertex-level contact annotation rather than boxes.

GRAB supplies mocap-quality hand and body meshes for natural grasps of everyday objects; HOI4D pushes to 2.4 million RGB-D frames with 4D hand-object models and category-level pose for 800 object instances across 16 categories[6]. Contact ground truth is scarcer than pose: ContactPose captured contact maps for everyday-object grasps using thermal cameras that read residual heat at contact points, and OakInk ties grasp type to intent, a precision pinch for small-part assembly versus a power grasp for a tool.

Teleoperation Data for Imitation Learning

Teleoperation datasets log the state-action trajectories behavior cloning needs: robot joint positions, gripper state, and task-success labels that pure vision data lacks. BridgeData V2 holds 60,000 trajectories over 24 skills on a WidowX 250 arm[7], DROID scales to 76,000 on bimanual Franka Panda arms across homes, offices, and labs, and ALOHA adds 1,000 bimanual mobile demonstrations for whole-body tasks like opening a drawer while steadying an object.

The interface sets the fidelity. Joystick control is cheap and low-bandwidth, mocap gloves are high-fidelity and expensive, and UMI splits the difference with a handheld gripper that carries its own cameras, letting operators demonstrate in the wild with no worn sensors. Viewpoint matters as much as the interface: when a robot wrist camera matches the operator head-mounted view, the sim-to-real gap shrinks, which is why RT-1 collected its 130,000 episodes with head-mounted cameras. Open X-Embodiment then aggregated roughly 1 million trajectories from 22 embodiments to benchmark cross-embodiment transfer[8].

3D Hand Pose Estimation and Parametric Models

3D hand pose is the substrate for grasp synthesis and contact prediction, and it ships in two representations: a 21-joint kinematic skeleton or a parametric mesh. MANO is the standard mesh, encoding articulation in 45 pose parameters over a low-dimensional basis, which is what lets a network regress a full hand surface instead of sparse joints. FreiHAND supplies 130,000 images with MANO fits for cross-shape generalization, InterHand2.6M adds 2.6 million frames of two-hand interaction for bimanual work, and DexYCB gives magnetic-tracker pose at 30 Hz.

Contact is the harder label. Thermal imaging (ContactPose) reads heat transfer but needs special cameras, embedded pressure sensors give true forces but cap object diversity, and TACTO renders vision-based tactile sensors (GelSight) down to sub-millimeter contact geometry, though few released datasets ship it. For storage, LeRobot keeps MANO parameters, joint positions, and contact masks as separate HDF5 datasets per episode[9], which is what makes random access and cross-dataset merging cheap.

Annotation Workflows and Quality Control

HOI annotation is a bootstrap-then-correct pipeline. Off-the-shelf detectors (MediaPipe, FrankMocap) seed hand pose, annotators fix joints in 3D against multi-view consistency, and object pose starts from CAD alignment or a category-level estimator before manual refinement. CVAT handles 3D cuboids and skeletons with keyframe interpolation, Labelbox runs three-annotator consensus with expert tie-breaks, and Scale AI's physical-AI service provides managed hand-pose and contact annotation at production scale.

Contact masks are the cost center: painting vertex-level contact on a hand mesh takes minutes of skilled work per grasp on a set like ContactPose, though Segments.ai SAM-based propagation cuts binary-mask work substantially. Judge a vendor on three numbers: joint-position error (Euclidean distance to ground truth), contact precision and recall at the vertex level, and downstream task-success rate once a policy trains on the data. Truelabel's data provenance system records per-trajectory metadata and consent artifacts, so you can vet annotation fidelity on a sample packet before committing to scale.

Simulation vs. Real-World HOI Data

Simulators like Isaac Gym and MuJoCo give infinite grasps and perfect ground truth but miss real contact physics, deformation, and appearance; real capture nails contact dynamics but costs infrastructure and manual labels. Benchmarks like DexArt generate synthetic articulated-object grasps (doors, drawers, scissors) with exact contact forces from rigid-body physics, yet a policy trained purely in simulation degrades on real articulated manipulation and recovers only after fine-tuning on real demonstrations. Domain randomization (varying texture, lighting, and camera pose in simulation) narrows that gap but never closes the contact-uncertainty part.

RoboNet runs the opposite play, using 15 million real frames from 7 platforms to calibrate simulator friction and compliance[10], while CALVIN ships paired real and simulated demonstrations so you can measure sim-to-real transfer loss directly.

Licensing and Commercial Use Rights

License terms are what usually block a commercial training run, more often than data quality itself. EPIC-KITCHENS annotations are MIT-licensed, but the underlying video still needs participant consent under GDPR Article 7. DexYCB is CC BY 4.0, so commercial use is fine with attribution. GRAB is CC BY-NC 4.0, so commercial use requires a separate negotiation. Teleoperation sets inherit their simulator terms: RoboNet's real subset is BSD-licensed, but its simulated episodes ride on MuJoCo assets under a non-commercial research license, and Open X-Embodiment bundles mixed licenses you must clear component by component.

Many robotics datasets on Hugging Face carry no explicit commercial-use clause, so the risk is silent[11]; budget months, not days, to clear a multi-source mix for foundation-model training. Truelabel's marketplace profiles licensing on each dataset card and delivers rights-cleared data with contributor consent artifacts and per-trajectory provenance.

DatasetLicenseCommercial training?
DexYCBCC BY 4.0Yes, with attribution
EPIC-KITCHENS (annotations)MITYes; video needs separate consent
GRABCC BY-NC 4.0No, negotiate first
RoboNet (real subset)BSDYes
Open X-EmbodimentMixedClear each component
Commercial-use posture of common HOI datasets

Procurement Checklist for HOI Datasets

What separates useful HOI datasets is diversity of objects and tasks, more than sheer volume. Open X-Embodiment aggregated 1 million trajectories across 22 embodiments and many task domains[8] precisely because breadth transfers better than repetition, so prioritize datasets with 100 or more object instances and 10 or more task categories over single-task sets with millions of repetitions.

Two format gotchas trip up buyers. For egocentric data, confirm the camera intrinsics (focal length, distortion coefficients) match your robot onboard camera. For teleoperation, verify the gripper-state encoding, binary open/close versus continuous position. RLDS standardizes these metadata fields so cross-dataset comparison is mechanical. Truelabel ships a sample packet with QA evidence before any scale commitment, so you validate hand-pose accuracy, contact density, and license terms on a pilot batch first.

  1. 01

    3D hand pose accuracy

    Joint-position error under 10 mm for grasp-synthesis work; ask for the error distribution, not just the mean.

  2. 02

    Contact annotation density

    Vertex-level contact maps on at least 20% of frames, consistent across adjacent frames.

  3. 03

    Object diversity

    50 or more object instances spanning rigid, articulated, and deformable categories.

  4. 04

    Task coverage

    Demonstrations of your target skills: pick-place, in-hand reorientation, tool use.

  5. 05

    Embodiment match

    End-effector geometry and DOF compatible with your deployment robot; teleoperation control frequency of 10 Hz or higher for contact-rich tasks.

  6. 06

    Licensing clarity

    Explicit commercial-use permission or a documented negotiation path, cleared per component for aggregated sets.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

External references and source context

  1. Project site

    DexYCB dataset statistics: 582,000 RGB-D frames with ground-truth 3D hand pose and object pose

    dex-ycb.github.io ↩
  2. Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100

    EPIC-KITCHENS-100 dataset paper: 100 hours, 700 action segments, 97 verbs, 300 nouns

    arXiv ↩
  3. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    DROID paper: 76,000 trajectories, 350 hours, 564 scenes, bimanual control

    arXiv ↩
  4. Ego4D: Around the World in 3,000 Hours of Egocentric Video

    Ego4D paper: large-scale annotated hand-object interactions across daily activities

    arXiv ↩
  5. Learning to Detect Human-Object Interactions

    HICO-DET statistics: 47,776 images, 150,000 human-object pairs, 600 interaction categories

    arXiv ↩
  6. HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction

    HOI4D project site: 2.4M RGB-D frames, 4D hand-object models, 800 object instances

    hoi4d.github.io ↩
  7. BridgeData V2: A Dataset for Robot Learning at Scale

    BridgeData V2 paper: 60,000 trajectories, 24 skills, WidowX 250 arm

    arXiv ↩
  8. Open X-Embodiment: Robotic Learning Datasets and RT-X Models

    Open X-Embodiment paper: 1M trajectories from 22 distinct embodiments

    arXiv ↩
  9. LeRobot documentation

    LeRobot documentation: HDF5 schema for hand pose, MANO parameters, contact masks

    Hugging Face ↩
  10. RoboNet: Large-Scale Multi-Robot Learning

    RoboNet paper: 15M frames from 7 robot platforms with simulation calibration

    arXiv ↩
  11. Robotics datasets on Hugging Face need a buyer-readiness layer

    Robotics datasets on Hugging Face frequently lack explicit commercial-use licensing

    Hugging Face ↩

More glossary terms

FAQ

What is the difference between hand pose estimation and hand-object interaction datasets?

Hand pose estimation datasets (e.g., FreiHAND, InterHand2.6M) focus solely on recovering 3D hand joint positions or mesh parameters from images, without object context. HOI datasets additionally capture object pose, contact surfaces, and task semantics (action labels, success flags). Hand pose datasets train pose estimators; HOI datasets train manipulation policies. DexYCB is an HOI dataset because it includes both hand pose and object pose with contact annotations. FreiHAND is a hand pose dataset because objects are present but not tracked or annotated.

Can I train a dexterous manipulation policy using only egocentric video datasets like EPIC-KITCHENS?

Egocentric video datasets provide rich semantic context (what objects are used, in what order) but typically lack 3D hand pose and end-effector trajectories required for imitation learning. You can train affordance models or task planners from egocentric video, then combine them with teleoperation datasets (DROID, BridgeData V2) that provide state-action pairs for low-level control. RT-2 demonstrates this approach, using web video for semantic grounding and robot teleoperation data for action execution. Pure egocentric video is insufficient for grasp synthesis without additional 3D pose annotation.

How do I verify contact annotation quality in an HOI dataset?

Request sample frames with contact masks overlaid on hand and object meshes. Check for: (1) vertex-level precision (contact labels at mesh resolution, not bounding-box approximations); (2) temporal consistency (contact regions should not flicker between adjacent frames); (3) physical plausibility (contact normals should oppose each other, indicating force closure). ContactPose provides thermal validation images showing heat transfer at contact points. For datasets without thermal validation, compute contact precision/recall against a held-out test set annotated by domain experts. Inter-annotator agreement (Dice coefficient) should exceed 0.85 for binary contact masks.

What is the typical cost to annotate 1,000 frames of HOI data with 3D hand pose and contact maps?

Annotation cost scales with fidelity. Skeletal hand pose (21 joints) with semi-automated tracking and manual correction is the cheapest tier; MANO mesh fitting with contact masks costs several times more per frame; full scene reconstruction with object pose and contact forces is the most expensive. Commercial annotation vendors price this work at a premium over academic teams, which use research assistants at lower hourly rates but longer timelines. Truelabel's collector network includes annotation specialists who price per-dataset based on task complexity, so buyers get a per-dataset quote against their fidelity and volume rather than a fixed per-frame rate.

Which HOI datasets support bimanual manipulation tasks?

Bimanual datasets include: GRAB (two-hand grasping of large objects), InterHand2.6M (two-hand interactions with 2.6M frames), Assembly101 (two-hand toy assembly with egocentric and exocentric views), ALOHA (1,000 bimanual mobile manipulation demonstrations), and DROID (bimanual Franka Panda teleoperation across 564 scenes). For policy training, verify that both hands are tracked simultaneously with synchronized timestamps. Some datasets (e.g., H2O) track hands independently, requiring post-processing to align left/right hand trajectories. ALOHA and DROID provide joint-space trajectories for both arms, simplifying behavior cloning.

How do I combine multiple HOI datasets with different annotation schemas for foundation model training?

Use RLDS (Reinforcement Learning Datasets) format to standardize heterogeneous datasets into a common schema with observation, action, reward, and metadata fields. LeRobot provides conversion scripts for 15+ robotics datasets, mapping dataset-specific fields (e.g., DexYCB's magnetic tracker data, EPIC-KITCHENS' verb-noun labels) into RLDS episodes. For hand pose, convert all representations to MANO parameters using off-the-shelf fitting tools (FrankMocap, HARP). For contact, binarize vertex-level masks into gripper-state proxies (contact detected → gripper closed). Truelabel's dataset cards expose schema mappings, enabling automated RLDS conversion for 80% of indexed HOI datasets.

Find datasets covering hand-object interaction

Truelabel surfaces vetted datasets and capture partners working with hand-object interaction. Send the modality, scale, and rights you need and we route you to the closest match.

Browse HOI datasets on truelabel