Physical-world AI data
Embodied AI Datasets
An embodied AI dataset contains observations, actions, scenes, tasks, or demonstrations that help models perceive and act in physical environments. Egocentric and first-person video can be part of embodied AI data when the viewpoint helps represent interaction, manipulation, navigation, or task context.
Quick facts
- Ego4D embodied benchmark context
- Ego4D contributes about 3,670 hours of egocentric video across 74 locations and 9 countries for benchmark and task-language planning.
- Ego-Exo4D skilled-activity context
- Ego-Exo4D contributes about 1,286 hours from 740 participants or camera wearers across 13 cities and 123 sites for paired-view skilled activities.
- EPIC-KITCHENS-100 action context
- EPIC-KITCHENS-100 contributes 90K action segments, 97 verb classes, and 300 noun classes for kitchen-activity framing under non-commercial public terms.
Comparison
| Perspective | Modality | Task | Domain |
|---|---|---|---|
| Egocentric | RGB, audio, IMU, gaze | Action anticipation | Human activity and tools |
| Robot-mounted | RGB-D, proprioception, actions | Manipulation | Warehouse, lab, home |
| Ego-exo | First- and third-person views | Skilled activity analysis | Sports, craft, repair |
How embodied AI data differs from generic video data
Generic video data may show scenes passively. Embodied AI datasets must map observations to actions, tasks, objects, environments, and sometimes robot state. Egocentric video is valuable when the actor viewpoint reveals task flow that an external camera may miss [1].
Modalities and annotations
Depending on the model, embodied AI data may require RGB video, depth, pose, action labels, narration, object state, robot proprioception, trajectories, success/failure labels, and source documentation. Ego-Exo4D shows why perspective can be part of the dataset design, not just a visual style [2].
Use cases in robotics and embodied AI
Embodied AI datasets can support manipulation, navigation, human demonstration learning, hand-object interaction analysis, tool-use reasoning, success/failure evaluation, and environment-specific model testing. The useful data shape changes by task: manipulation may need object state and action labels, while navigation may need scene context, trajectory, and success criteria.
Public datasets may not match deployment needs
Public datasets help benchmark model behavior and define terms. They may not match the buyer's target embodiment, environment, object set, rights posture, consent basis, or QA rubric. Those gaps belong in a custom collection plan rather than in unsupported performance claims.
Embodied AI dataset taxonomy
Embodied AI data connects observations to physical actions, tasks, states, environments, and outcomes. Generic video may show a scene; embodied datasets need fields that let a model or evaluator reason about what the agent did and whether it succeeded, plus provenance and rights artifacts for buyer review before production use.
| Type | Required fields | Best use |
|---|---|---|
| Egocentric human video | Viewpoint, task labels, consent, annotations | Perception, task grammar, HOI context |
| Robot demonstrations | Observations, state, actions, outcomes | Imitation learning and VLA fine-tuning |
| Teleoperation trajectories | Operator control, action stream, robot state, timestamps | Policy learning and recovery data |
| Eval data | Held-out scenes, scoring rubric, failures | Supplier QA and regression testing |
| Simulation/synthetic | Task definitions, environment config, domain gap notes | Ablations and coverage planning |
Public examples and what they do not prove
Public embodied-AI references answer different questions. DROID and Open X-Embodiment guide robot-trajectory schema and embodiment diversity; Ego4D and EPIC-KITCHENS guide first-person task understanding; LeRobot-style Hub datasets guide packaging and training workflow. None of those examples automatically proves commercial-use rights, target-domain fit, or supplier readiness for your robot.
| Reference type | Use it for | Replace or verify before production |
|---|---|---|
| DROID / robot demos | trajectory fields and scene diversity | target robot, task, rights, and eval split |
| Open X-Embodiment | multi-embodiment schema expectations | your embodiment and action representation |
| Ego4D / EPIC-KITCHENS | egocentric task vocabulary | commercial rights and deployment domain |
| LeRobot datasets | format and tooling | load test, license, provenance, and field coverage |
Minimum schema for embodied AI data
Request observation streams, timestamps, action representation, robot or body state, language instruction when relevant, task boundary, object state, success/failure labels, environment metadata, consent/provenance, license terms, and delivery format. If the model is a VLA training stack, require language-aligned action/state episodes; if the program is commercial, route rights and provenance through egocentric data licensing and the robot training marketplace.
Train, eval, simulation, and teleoperation are separate asks
Training data should cover the target distribution; eval data should be held out and failure-heavy; simulation data should state the domain gap; teleoperation data should preserve controller provenance and actions. One vendor sample can include all four only if the manifest names which rows are train, eval, sim-derived, or teleoperated and why each belongs there.
Embodied AI buyer checklist
Before procurement, ask four questions. What signal teaches the policy: observation, action, state, language, or contact? What split measures success without leakage? What public reference only informs schema rather than rights? What artifact proves consent, provenance, and format loadability? The answer decides whether the request belongs in VLA training data, teleoperation, eval data, egocentric video, or the robotics data marketplace.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
The Ego4D paper is the source-backed reference for first-person daily-life activity video and benchmark design.
arXiv ↩ - Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
The Ego-Exo4D paper describes skilled human activity from first- and third-person perspectives.
arXiv ↩ - Egocentric video remains useful but incomplete for robot data buyers
Ego4D is an official public reference for egocentric video dataset scope, access, and dataset documentation.
ego4d-data.org - Ego-Exo4D project site
Ego-Exo4D is the official project source for paired first-person and third-person skilled-activity capture.
ego-exo4d-data.org - Ego-Exo4D annotations documentation
Ego-Exo4D annotation documentation supports dataset-structure and skilled-activity-label discussion.
docs.ego-exo4d-data.org - EPIC-KITCHENS project site
EPIC-KITCHENS is an official project reference for egocentric kitchen-activity data.
epic-kitchens.github.io - Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100
The EPIC-KITCHENS-100 paper supports public kitchen-activity benchmark facts and caveats.
arXiv - EPIC-KITCHENS-100 annotations license
The EPIC-KITCHENS-100 annotation license is a visible source for non-commercial licensing caveats.
GitHub - truelabel egocentric data glossary
Internal contextual link to the egocentric data definition.
truelabel.ai - truelabel sourcing brief intake
Internal contextual link to Truelabel's sourcing intake workflow.
truelabel.ai - truelabel egocentric warehouse video sourcing spec
Internal contextual link to warehouse egocentric video sourcing.
truelabel.ai - truelabel egocentric kitchen video sourcing spec
Internal contextual link to kitchen egocentric video sourcing.
truelabel.ai - truelabel industrial egocentric video sourcing spec
Internal contextual link to industrial egocentric video sourcing.
truelabel.ai - truelabel warehouse robotics data sourcing
Internal contextual link to warehouse robotics data sourcing.
truelabel.ai - truelabel kitchen manipulation data sourcing
Internal contextual link to kitchen manipulation data sourcing.
truelabel.ai - truelabel LeRobot format guide
Internal contextual link to the LeRobot format guide.
truelabel.ai - truelabel LeRobot dataset alternative comparison
Internal contextual link to the LeRobot dataset alternative comparison.
truelabel.ai - truelabel eval data for robotics hub
Internal contextual link to robotics eval data sourcing.
truelabel.ai - truelabel teleoperation training-data page
Internal contextual link to teleoperation training data sourcing.
truelabel.ai - truelabel robot demonstrations training-data page
Internal contextual link to robot demonstration training data sourcing.
truelabel.ai - truelabel hand-object interaction data page
Internal contextual link to hand-object interaction training data requirements.
truelabel.ai - truelabel egocentric video datasets hub
Internal contextual link to the egocentric video datasets hub.
truelabel.ai
More glossary terms
FAQ
What is an embodied AI dataset?
It is data describing agents acting in physical environments, often including observations, actions, tasks, scenes, and annotations.
How is egocentric video useful for embodied AI?
It can show task execution from the actor viewpoint, including hands, objects, occlusion, tool use, and sequential context.
What data does a robotics model need?
It depends on the task, but common requirements include observations, actions, task labels, object state, metadata, and quality checks tied to the deployment context.
What is the difference between embodied AI data and generic video data?
Embodied AI data is connected to action and physical-world tasks; generic video may lack action labels, embodiment context, rights details, and task-specific metadata.
How is an embodied AI dataset different from a video dataset?
Generic video may show scenes; embodied AI data needs action/task labels, embodiment context, timestamps, object/environment state, and often robot or human demonstration metadata.
Which embodied dataset type do I need for VLA training?
Use VLA-shaped episodes with observations, language instructions, action/state streams, success labels, embodiment metadata, and a consistent format such as LeRobot or RLDS.
Find datasets covering embodied AI datasets
Truelabel surfaces vetted datasets and capture partners working with embodied AI datasets. Send the modality, scale, and rights you need and we route you to the closest match.
Map your robotics data requirement to capture, consent, and QA constraints