Dataset hub
Egocentric Video Datasets
An egocentric video dataset contains video captured from the viewpoint of a person, wearable device, or robot. Robotics and physical AI teams use these datasets to study hands, objects, tools, spaces, task progress, and interaction context, while separately evaluating consent, provenance, licensing, and task fit.
Comparison
| Public example | Useful for | Caveat to verify |
|---|---|---|
| Ego4D | Broad egocentric daily-life research | Access terms and task fit |
| Ego-Exo4D | Paired first- and third-person skilled activities | Annotation and viewpoint fit |
| EPIC-KITCHENS | Egocentric kitchen activity recognition | Non-commercial licensing caveats |
What egocentric datasets typically capture
Egocentric video datasets usually capture a first-person view of task execution. Public references such as Ego4D, Ego-Exo4D, and EPIC-KITCHENS help teams benchmark tasks and terminology, but each source must be evaluated for access terms, scope, modality, and limitations before use [1] [2] [3].
Modalities and tasks to specify
A buyer brief should separate modality from task. Modality describes what is captured; task describes what the model should learn or evaluate.
| Dimension | Examples | Why it matters |
|---|---|---|
| Modality | RGB, audio, gaze, IMU, depth, pose, narration | Defines capture hardware and annotation needs |
| Task | Action recognition, anticipation, hand-object interaction, SLAM | Defines labels, clips, and acceptance checks |
| Domain | Kitchen, workshop, warehouse, home, retail | Defines object set and environment context |
Public datasets are references, not commercial supply by default
Public datasets are useful for benchmarking and task language. They should not be treated as commercial training supply unless the source terms, consent basis, and downstream license allow that use. EPIC-KITCHENS materials, for example, carry visible non-commercial licensing language in the public annotation license [4].
Dataset selection matrix for robotics buyers
Choose a dataset by task fit, viewpoint, annotation depth, access path, and rights posture rather than by name recognition. Public egocentric datasets are strongest as benchmark references and schema inspiration; custom sourcing is the safer route when commercial rights, target-domain coverage, or consent artifacts matter.
| Reference | Best public use | Common commercial gap |
|---|---|---|
| Ego4D | Broad first-person daily-life research and task vocabulary | May not match target facility, object set, or commercial-use terms |
| Ego-Exo4D | Paired first-/third-person skilled activity and multimodal annotation design | May be richer than needed for single-view deployment capture |
| EPIC-KITCHENS | Kitchen verb-noun action structure and benchmark framing | Public materials carry non-commercial caveats |
| HOI / hand datasets | Hand-object labels, pose, contact, and manipulation taxonomy | Need official source/license review before commercial use |
| Custom Truelabel capture | Buyer-specific domain, hardware, consent, QA, and delivery format | Requires a clear spec and pilot review before scale |
Which egocentric dataset fits which robotics task?
Hand-object interaction pages need contact, object state, grasp phase, and failure labels; the egocentric data definition is only the starting point. Kitchen pages need verb/noun/action boundaries and private-space consent, which is why kitchen capture should route through a kitchen egocentric video spec. Warehouse pages need SKU, scan, bin, shift, and facility-permission metadata; compare the warehouse egocentric video brief before requesting volume. VLA or robot policy work usually needs synchronized robot actions/states in addition to first-person video.
Minimum egocentric dataset specification
A production-ready egocentric brief should name the camera viewpoint, device, FPS/resolution, audio policy, sensor streams, task taxonomy, environment coverage, participant and bystander procedure, annotation schema, QA acceptance rules, license/model-use rights, and target delivery format such as MP4+JSON, LeRobot, RLDS, HDF5, or MCAP when robot state exists. Run the rights review through egocentric data licensing before a public benchmark becomes a training source, then move the approved collection into a robot training data marketplace request with a pilot gate. The brief should also state what is deliberately out of scope: sensitive rooms, minors, screens, badges, audio, hazardous equipment, public bystanders, or unsupported robot-action claims. Those exclusions are content quality, not bureaucracy, because they prevent unusable clips from entering the training queue.
Custom collection workflow
The safest workflow is scope, sample, audit, then scale. Start with a public-reference table, translate the target task into a supplier-ready sourcing brief, ask for 50-200 pilot clips or episodes, inspect accepted and rejected samples, verify consent/provenance artifacts, and only then fund the full collection. This prevents a dataset from becoming a large pile of unlicensed or untrainable footage. A good pilot includes a manifest, raw clips, redaction log, task labels, reviewer notes, and a list of failure cases that will be represented in the full batch. If the buyer cannot reject the pilot with a written reason, the full program is not specified tightly enough.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- Egocentric video remains useful but incomplete for robot data buyers
Ego4D is an official public reference for egocentric video dataset scope, access, and dataset documentation.
ego4d-data.org ↩ - Ego-Exo4D project site
Ego-Exo4D is the official project source for paired first-person and third-person skilled-activity capture.
ego-exo4d-data.org ↩ - EPIC-KITCHENS project site
EPIC-KITCHENS is an official project reference for egocentric kitchen-activity data.
epic-kitchens.github.io ↩ - EPIC-KITCHENS-100 annotations license
The EPIC-KITCHENS-100 annotation license is a visible source for non-commercial licensing caveats.
GitHub ↩ - Ego4D: Around the World in 3,000 Hours of Egocentric Video
The Ego4D paper is the source-backed reference for first-person daily-life activity video and benchmark design.
arXiv - Ego-Exo4D annotations documentation
Ego-Exo4D annotation documentation supports dataset-structure and skilled-activity-label discussion.
docs.ego-exo4d-data.org - Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
The Ego-Exo4D paper describes skilled human activity from first- and third-person perspectives.
arXiv - Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100
The EPIC-KITCHENS-100 paper supports public kitchen-activity benchmark facts and caveats.
arXiv - truelabel industrial egocentric video sourcing spec
Internal contextual link to industrial egocentric video sourcing.
truelabel.ai - truelabel VLA training data sourcing
Internal contextual link to VLA training data sourcing.
truelabel.ai - truelabel warehouse robotics data sourcing
Internal contextual link to warehouse robotics data sourcing.
truelabel.ai - truelabel kitchen manipulation data sourcing
Internal contextual link to kitchen manipulation data sourcing.
truelabel.ai - truelabel LeRobot format guide
Internal contextual link to the LeRobot format guide.
truelabel.ai - truelabel LeRobot dataset alternative comparison
Internal contextual link to the LeRobot dataset alternative comparison.
truelabel.ai - truelabel eval data for robotics hub
Internal contextual link to robotics eval data sourcing.
truelabel.ai - truelabel teleoperation training-data page
Internal contextual link to teleoperation training data sourcing.
truelabel.ai - truelabel robot demonstrations training-data page
Internal contextual link to robot demonstration training data sourcing.
truelabel.ai - truelabel hand-object interaction data page
Internal contextual link to hand-object interaction training data requirements.
truelabel.ai - truelabel egocentric video datasets hub
Internal contextual link to the egocentric video datasets hub.
truelabel.ai
FAQ
What is an egocentric video dataset?
It is a dataset of videos recorded from the actor or agent viewpoint, often using wearable, head-mounted, handheld, or robot-mounted cameras.
What are egocentric video datasets used for?
They support research and planning for action recognition, anticipation, hand-object interaction, task understanding, and embodied AI evaluation.
What modalities are common in egocentric video data?
Common modalities include RGB video, audio, gaze, IMU, depth, pose, point cloud data, narration, and task annotations.
Can public egocentric datasets be used commercially?
Do not assume so. Review the dataset source, license, consent basis, and downstream model-use terms before using any public dataset in a commercial program.
What is the best egocentric video dataset for robotics?
There is no universal best. Ego4D is broad, Ego-Exo4D helps with paired views, EPIC-KITCHENS is focused on kitchens, and HOI-oriented datasets may fit manipulation better. Choose by task, modality, labels, and verified rights.
What fields should an egocentric dataset include for robot learning?
At minimum: synchronized video, task labels, timestamps, object or hand annotations when needed, episode boundaries, success/failure labels, consent/provenance artifacts, license terms, and delivery manifest. For policy learning, add robot state and action streams.
When is custom capture better than a public egocentric benchmark?
Use custom capture when public data misses the target environment, object mix, device, region, privacy artifacts, format, or commercial model-use rights.
Looking for egocentric video datasets?
Specify modality, task, environment, requested rights posture, and delivery format. Truelabel routes the request to candidate capture partners and helps scope consent/provenance artifacts and commercial licensing requirements for buyer review before delivery.
Compare public dataset limits with a custom collection plan