Manipulation data
Hand-Object Interaction Data for Robotics
Hand-object interaction data captures how hands approach, grasp, move, use, and release objects. For robotics, egocentric video can help document interaction context, but useful datasets also need task boundaries, object state, annotations, quality controls, and governance review.
Comparison
| Annotation | Model use | Quality failure to catch |
|---|---|---|
| Hand visibility | Grasp and approach context | Hands leave frame or occlude object |
| Object state | Before/after task understanding | State changes are missing or ambiguous |
| Action phase | Temporal segmentation | Task boundaries are inconsistent |
| Failure label | Robustness and eval design | Only clean successes are captured |
Why egocentric video is useful but not sufficient
Egocentric video can show hands, tools, object contact, and occlusion from the actor's perspective. It does not by itself solve manipulation: buyers still need annotations, object-state definitions, task boundaries, metadata, and a QA rubric [1].
Public dataset examples and limitations
Ego4D, Ego-Exo4D, and EPIC-KITCHENS can help frame hand-object interaction tasks, skilled activities, and kitchen activity recognition. They should be treated as public research references unless the intended use, access terms, consent posture, and license are separately verified [2] [3] [4].
Data quality and failure modes
A robotics collection plan should include failure cases, not only successful demonstrations. Watch for unstable cameras, missing object state, unlabelled task interruptions, poor lighting, incomplete consent documentation, and samples that cannot be aligned to the model's input format.
HOI task taxonomy for robot learning
Useful HOI data describes the sequence around contact, not just a frame with a hand and object. Capture approach, pre-grasp, contact, grasp, manipulate, tool use, bimanual coordination, state change, release/place, and failure or recovery phases.
| Phase | Labels to request | Failure to catch |
|---|---|---|
| Approach/contact | hand visibility, object identity, contact state | hands out of frame or ambiguous object |
| Grasp/manipulate | grasp type, action phase, object state | occlusion at the critical contact moment |
| Tool use | tool class, material, action boundary | tool/object labels do not match task taxonomy |
| Release/failure | success, slip, drop, recovery | only clean successes are captured |
Minimum annotation schema
Request hand visibility, 2D/3D hand pose where available, object identity/category, object pose/state when needed, contact state, grasp type, action phase, tool class, occlusion flag, success/failure, timestamps, consent/provenance artifacts, and target delivery format.
HOI sample manifest
A hand-object interaction sample manifest should include clip_id, task phase, camera view, hand visibility score, object ID/category/material, contact state, grasp type, object state before/after, occlusion flag, success/failure label, rejected reason, consent/provenance artifact, and delivery format. Link the manifest to adjacent warehouse, kitchen, or industrial capture constraints when the setting changes the privacy or object taxonomy.
Accepted and rejected HOI clips
Accepted clips show the decisive contact moment, the object state change, and enough context to understand the intended task. Rejected clips are equally valuable: hands out of frame, reflective objects, hidden contact, unlabeled tool change, ambiguous success, missing consent, or footage without a task boundary should all be preserved as rejection examples before scale-up.
Task-fit matrix for HOI buyers
Grasping, tool use, bimanual handoff, deformables, and kitchen manipulation need different HOI fields. Grasping needs contact outcome; tool use needs tool state and force proxy; bimanual work needs left/right synchronization; deformables need shape state before and after; kitchen tasks need object state plus privacy artifacts. Route the adjacent procurement page accordingly: grasping, dexterous manipulation, bimanual manipulation, workshop capture, or kitchen sourcing.
| Task | Must capture | Reject if |
|---|---|---|
| Grasping | approach, contact, slip/drop | only final success visible |
| Tool use | tool class, material, action boundary | tool/object mismatch |
| Bimanual handoff | both hands, sync, transfer state | one side out of frame |
| Deformables | shape before/after, force proxy | state change unlabeled |
| Kitchen | verb-noun, appliance/object state | privacy artifacts missing |
Public HOI references and commercial gaps
Public egocentric and HOI datasets can inform task vocabulary and annotation design. They should not be treated as commercial training supply until official license, access, consent, and model-use terms are reviewed through licensing/provenance review. For production manipulation, request a pilot through the robot training data marketplace so contact labels, failure cases, and rights are audited together.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
The Ego4D paper is the source-backed reference for first-person daily-life activity video and benchmark design.
arXiv ↩ - Egocentric video remains useful but incomplete for robot data buyers
Ego4D is an official public reference for egocentric video dataset scope, access, and dataset documentation.
ego4d-data.org ↩ - Ego-Exo4D project site
Ego-Exo4D is the official project source for paired first-person and third-person skilled-activity capture.
ego-exo4d-data.org ↩ - EPIC-KITCHENS-100 annotations license
The EPIC-KITCHENS-100 annotation license is a visible source for non-commercial licensing caveats.
GitHub ↩ - Ego-Exo4D annotations documentation
Ego-Exo4D annotation documentation supports dataset-structure and skilled-activity-label discussion.
docs.ego-exo4d-data.org - Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
The Ego-Exo4D paper describes skilled human activity from first- and third-person perspectives.
arXiv - EPIC-KITCHENS project site
EPIC-KITCHENS is an official project reference for egocentric kitchen-activity data.
epic-kitchens.github.io - Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100
The EPIC-KITCHENS-100 paper supports public kitchen-activity benchmark facts and caveats.
arXiv - truelabel egocentric data glossary
Internal contextual link to the egocentric data definition.
truelabel.ai - truelabel sourcing brief intake
Internal contextual link to Truelabel's sourcing intake workflow.
truelabel.ai - truelabel VLA training data sourcing
Internal contextual link to VLA training data sourcing.
truelabel.ai - truelabel warehouse robotics data sourcing
Internal contextual link to warehouse robotics data sourcing.
truelabel.ai - truelabel kitchen manipulation data sourcing
Internal contextual link to kitchen manipulation data sourcing.
truelabel.ai - truelabel LeRobot format guide
Internal contextual link to the LeRobot format guide.
truelabel.ai - truelabel LeRobot dataset alternative comparison
Internal contextual link to the LeRobot dataset alternative comparison.
truelabel.ai - truelabel eval data for robotics hub
Internal contextual link to robotics eval data sourcing.
truelabel.ai - truelabel teleoperation training-data page
Internal contextual link to teleoperation training data sourcing.
truelabel.ai - truelabel robot demonstrations training-data page
Internal contextual link to robot demonstration training data sourcing.
truelabel.ai - truelabel hand-object interaction data page
Internal contextual link to hand-object interaction training data requirements.
truelabel.ai - truelabel egocentric video datasets hub
Internal contextual link to the egocentric video datasets hub.
truelabel.ai
More glossary terms
FAQ
What is hand-object interaction data?
It is data that captures hands interacting with objects, including approach, grasp, manipulation, tool use, release, and task outcome.
Why is egocentric video useful for hand-object interaction?
The actor viewpoint often keeps hands, tools, and object contact in the frame, which can make interaction context easier to inspect.
What annotations are needed for hand-object interaction datasets?
Common annotations include action phase, object identity, object state, hand visibility, grasp type, success or failure, and task boundaries.
What are common quality failures in robotics manipulation data?
Common failures include occluded objects, inconsistent task boundaries, missing failures, motion blur, poor synchronization, and incomplete metadata or consent evidence.
Which annotations matter most in an HOI dataset?
Start with object identity, hand visibility, contact state, grasp type, action phase, object state before/after, success/failure, occlusion, and timestamps. Add pose, depth, tactile, or robot state only when the model consumes those fields.
How is HOI data different from generic manipulation data?
Generic manipulation data may record robot actions or outcomes; HOI data focuses on fine-grained hand/end-effector and object interaction: approach, contact, grasp, tool use, object state, occlusion, and recovery.
Find datasets covering hand-object interaction data
Truelabel surfaces vetted datasets and capture partners working with hand-object interaction data. Send the modality, scale, and rights you need and we route you to the closest match.
Map your robotics data requirement to capture, consent, and QA constraints