Capture method
Wearable Camera Datasets
Wearable camera datasets are first-person video or sensor collections captured from cameras worn by people, such as head-mounted action cameras, body cameras, AR glasses, or chest rigs. For robotics and physical AI, they are useful only when the footage is paired with task metadata, consent/provenance artifacts, annotation rules, and license terms that match the intended model use.
Quick facts
- Ego4D wearable-scale reference
- Ego4D is a source-backed egocentric reference with about 3,670 hours from 74 locations and 9 countries; review access terms before use.
- Ego-Exo4D capture breadth
- Ego-Exo4D includes skilled activities from 740 participants or camera wearers across 13 cities and 123 sites, using paired ego/exo viewpoints.
- Privacy planning fields
- Wearable-camera briefs should evaluate faces, voices, screens, homes, workplaces, locations, bystanders, consent, retention, and security before capture.
Comparison
| Device type | Strength | Limitation and privacy risk |
|---|---|---|
| Head-mounted camera | Stable actor viewpoint | Faces, homes, screens, and bystanders may enter frame |
| Chest or body camera | Longer wearability | Hands and small objects may be occluded |
| Glasses-style device | Natural hands-free viewpoint | Device policies and privacy expectations need review |
| Action camera | Rugged capture | Motion blur and comfort can affect QA |
Relationship to egocentric and first-person data
Wearable cameras are a capture method. Egocentric or first-person data is the viewpoint category. A wearable camera can produce useful egocentric video, but the dataset still needs task metadata, source documentation, consent review, and quality checks [1].
Capture devices and tradeoffs
Head-mounted cameras, body cameras, glasses-style devices, and rugged action cameras can all produce wearable-camera datasets. The right device depends on task duration, hand visibility, comfort, field of view, motion blur tolerance, audio policy, and whether the environment may expose faces, badges, screens, or private spaces.
Dataset fields and annotations
A wearable-camera brief should state capture rig, frame rate, field of view, audio policy, task boundary, object list, annotation type, bystander rules, and accepted failure modes. Ego-Exo4D annotation documentation is a useful public reference for thinking about skilled-activity labels [2].
Common use cases
Wearable-camera datasets are most useful when a team needs the actor's task context: hand-object interaction, tool use, assembly, repair, navigation through a workspace, safety review, or first-person demonstrations for robotics evaluation. They are less useful when the model needs a fixed external camera angle or robot proprioception as the primary signal.
Privacy and bystander risks
Wearable capture can enter homes, workplaces, screens, badges, faces, voices, and locations. Buyers should plan consent, notice, retention, de-identification review, and provider questions before capture begins.
Wearable camera datasets vs egocentric video datasets
Wearable camera describes how the data is captured; egocentric describes the viewpoint. A head-mounted camera, smart-glasses rig, chest camera, wrist camera, or robot-mounted camera can all produce first-person context, but they create different hand visibility, motion blur, privacy, metadata, and QA requirements.
| Term | Means | Buyer implication |
|---|---|---|
| Wearable camera dataset | Captured by a device worn by a person | Specify rig, comfort, FOV, audio, timestamps, and consent artifacts |
| Egocentric video dataset | Captured from the actor or agent viewpoint | Can include wearable, handheld, wrist, or robot-mounted views |
| First-person / POV video | Plain-language viewpoint term | Raw POV footage is not training-ready without metadata and rights |
| Robot wrist or gripper camera | Agent-adjacent camera on the robot | Needs action/state sync if used for policy learning |
Minimum buyer spec for wearable-camera capture
A useful request is more than 'wearable video'. Buyers should specify device type, mounting position, field of view, frame rate, audio policy, camera intrinsics when needed, task boundaries, environment coverage, object taxonomy, bystander procedure, consent evidence, retention/deletion policy, and delivery format. Ask for accepted and rejected sample clips before scaling.
- Capture: head/chest/glasses/wrist rig, FOV, FPS, resolution, stabilization, audio, IMU/gaze if available.
- Dataset: task labels, clip boundaries, object list, location class, wearer/session IDs, camera metadata, and QA rejection reasons.
- Governance: contributor consent, bystander handling, site permission, redaction log, license/model-use rights, and provenance chain.
Public benchmarks vs custom wearable capture
Ego4D, Ego-Exo4D, EPIC-KITCHENS, Project Aria-style resources, and related public datasets can help teams understand task vocabulary and capture tradeoffs. They should not be treated as commercial supply unless official access, consent, and license terms support the intended use. Custom capture is usually needed when the buyer needs a specific device, facility, region, object set, or commercial-rights package. Keep the official dataset page, license text, and access terms beside any citation so reviewers can separate public benchmark value from permitted production use.
Device and rig selection matrix
Choose the camera rig by task, not by brand. Head-mounted cameras capture gaze-adjacent action; chest rigs are more stable but lower; wrist or tool-mounted cameras expose contact but miss global context; AR glasses can add gaze, IMU, SLAM, and audio but create higher privacy review burden. Use egocentric data to define the viewpoint, then compare the setting-specific warehouse, kitchen, or industrial sourcing brief before choosing hardware. The rig choice also decides whether bystanders can see that capture is happening, so it belongs in the consent plan rather than in a late technical appendix.
| Rig | Best for | QA risk |
|---|---|---|
| Head-mounted | gaze-adjacent hands and tools | comfort and bystander visibility |
| Chest | stable long capture | hands may sit low in frame |
| Wrist/tool | contact and grasp detail | global context missing |
| AR glasses | gaze/IMU/SLAM-rich capture | privacy and access constraints |
Sample manifest for wearable-camera data
A useful sample manifest lists clip_id, wearer/session ID, device, mount, FOV, FPS/resolution, audio policy, task label, start/end time, environment, hands-visible flag, bystander/privacy handling, consent/provenance artifact, requested license terms, accepted/rejected reason, and delivery file paths. Ask for one accepted clip and one rejected clip per task family: a clean grasp, a hands-out-of-frame failure, a privacy rejection, and a label-ambiguity rejection. This turns the wearable-camera page from a hardware description into a QA rubric a buyer can send to suppliers, while leaving final rights and privacy determinations to buyer review.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- Egocentric video remains useful but incomplete for robot data buyers
Ego4D is an official public reference for egocentric video dataset scope, access, and dataset documentation.
ego4d-data.org ↩ - Ego-Exo4D annotations documentation
Ego-Exo4D annotation documentation supports dataset-structure and skilled-activity-label discussion.
docs.ego-exo4d-data.org ↩ - Ego4D: Around the World in 3,000 Hours of Egocentric Video
The Ego4D paper is the source-backed reference for first-person daily-life activity video and benchmark design.
arXiv - Ego-Exo4D project site
Ego-Exo4D is the official project source for paired first-person and third-person skilled-activity capture.
ego-exo4d-data.org - Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
The Ego-Exo4D paper describes skilled human activity from first- and third-person perspectives.
arXiv - EPIC-KITCHENS project site
EPIC-KITCHENS is an official project reference for egocentric kitchen-activity data.
epic-kitchens.github.io - truelabel physical AI data marketplace bounty intake
Internal contextual link to Truelabel's physical AI and robotics data marketplace.
truelabel.ai - truelabel egocentric data licensing hub
Internal contextual link to egocentric data licensing and provenance guidance.
truelabel.ai - truelabel sourcing brief intake
Internal contextual link to Truelabel's sourcing intake workflow.
truelabel.ai - truelabel VLA training data sourcing
Internal contextual link to VLA training data sourcing.
truelabel.ai - truelabel warehouse robotics data sourcing
Internal contextual link to warehouse robotics data sourcing.
truelabel.ai - truelabel kitchen manipulation data sourcing
Internal contextual link to kitchen manipulation data sourcing.
truelabel.ai - truelabel LeRobot format guide
Internal contextual link to the LeRobot format guide.
truelabel.ai - truelabel LeRobot dataset alternative comparison
Internal contextual link to the LeRobot dataset alternative comparison.
truelabel.ai - truelabel eval data for robotics hub
Internal contextual link to robotics eval data sourcing.
truelabel.ai - truelabel teleoperation training-data page
Internal contextual link to teleoperation training data sourcing.
truelabel.ai - truelabel robot demonstrations training-data page
Internal contextual link to robot demonstration training data sourcing.
truelabel.ai - truelabel hand-object interaction data page
Internal contextual link to hand-object interaction training data requirements.
truelabel.ai - truelabel egocentric video datasets hub
Internal contextual link to the egocentric video datasets hub.
truelabel.ai
More glossary terms
FAQ
What is a wearable camera dataset?
It is video or sensor data captured from cameras worn on the head, body, glasses, or similar devices.
What data can wearable cameras capture?
They can capture RGB video, audio, motion context, gaze-like viewpoint, hands, tools, object interactions, task sequence, and environment cues depending on the device.
How are wearable camera datasets used in robotics?
They can document human demonstrations and object interactions from the actor viewpoint for task analysis, imitation-learning planning, or evaluation design.
What privacy issues exist with wearable camera data?
Wearable cameras can capture identifiable people, voices, screens, homes, workplaces, locations, and bystanders who are not the primary contributor.
How is a wearable camera dataset different from an egocentric video dataset?
Wearable camera describes the capture method; egocentric describes the actor or agent viewpoint. Many egocentric datasets are wearable-camera datasets, but robot wrist, handheld POV, and gripper cameras can also be egocentric.
What fields should I request from a wearable-camera supplier?
Request device type, FOV, frame rate, audio policy, timestamps, task labels, clip boundaries, bystander handling, consent/provenance artifacts, license terms, delivery manifest, and QA rejection rules.
When should I collect custom wearable-camera data?
Collect custom data when public benchmarks do not cover your target environment, device, task taxonomy, privacy artifacts, commercial-rights posture, or delivery format.
Find datasets covering wearable camera datasets
Truelabel surfaces vetted datasets and capture partners working with wearable camera datasets. Send the modality, scale, and rights you need and we route you to the closest match.
Discuss a consented data collection brief