truelabelRequest dataEarnRequest

Buyer routing

Physical AI data marketplaces — when to use which

Six intake pages cover the most common physical AI buyer intents. The right one depends on modality (egocentric vs teleop vs robot-demos), output (VLA pair vs eval set), and scale (full capture vs validation sample). Use the decision table to route to the marketplace that matches your spec.

Decision rule

Pick the marketplace that matches your intent
If you need…Start atWhy
I need broad physical AI data — discovery and selection/physical-ai-data-marketplaceCross-modality marketplace covering teleop, egocentric, demonstrations, and eval.
First-person / wearable / hand-pose / consented egocentric/egocentric-data-licensingEgocentric capture with explicit licensing review and consent artifacts.
Robot teleoperation traces with synchronized state and action/teleoperation-data-marketplaceTeleop-specific intake — robot embodiment, action stream, success/failure flags.
Robot demonstrations and trajectories for policy training/robot-training-data-marketplaceGeneralist robotics intake covering manipulation, navigation, eval sets.
Vision-language-action paired data for OpenVLA / RT-2 / π0 / GR00T/vla-training-dataVLA-specific routing — paired observations + instructions + actions.
Smaller eval / validation / benchmark sample before scale/eval-data-for-roboticsEval-bound intake with scoped sample size and acceptance criteria.

All 9 marketplaces

Egocentric video data hub

Egocentric Video Data: Capture, License & Deliver for Physical AI

Egocentric video data is first-person video recorded from a head-mounted or wearable camera while a person performs a real-world task. It teaches physical-AI and robotics models hand-object interaction and viewpoint grounding. Truelabel routes buyer specs to candidate collectors for sample review and robotics-ready delivery with consent/provenance artifacts.

Buyer intake

Physical AI data marketplace

Truelabel is a physical AI data marketplace where candidate capture suppliers reviewed against the buyer spec deliver egocentric video, teleoperation traces, robot demonstrations, and evaluation datasets with requested commercial-use terms, contributor consent/provenance artifacts, and per-buyer fitness review. US search demand for 'physical AI' grew 3.5× between May 2025 and April 2026 (1,900 → 6,600 monthly searches), driven by humanoid programs (NVIDIA GR00T N1, Figure AI commercial deployments), open-source policy releases (Open X-Embodiment, RoboCasa), and the maturation of factory-style data pipelines that replace passive web scraping. Buyers post a sourcing request, review matched samples, and ingest data with rights and metadata attached.

First-person data

Egocentric data licensing for commercial AI training

Egocentric data is first-person video or sensor data captured from the perspective of a person performing real tasks — truelabel sources licensed egocentric footage benchmarked against public corpora like Ego4D's 3,670 hours of first-person daily-life activity. Buyers define environment, task, consent, and metadata requirements before suppliers submit samples for review.

Robot action data

Teleoperation data marketplace

A teleoperation data marketplace lets robotics teams source synchronized camera streams, joint states, end-effector poses, action traces, task labels, and success/failure metadata captured while a human operator controls the robot — and Truelabel benchmarks teleop sourcing against public references like RoboSet's 9,500 teleoperated trajectories. Truelabel matches teleop sourcing requests to candidate capture suppliers reviewed against the buyer spec and routes samples through buyer review before scale. The decision is a scenario, not a winner: custom teleop capture on your embodiment, a public baseline for research, a managed enterprise program, or tooling-plus-capture. Public datasets (DROID, BridgeData V2, RoboSet, AgiBot World) are baselines — they rarely satisfy your exact deployment distribution, consent, or embodiment. TrueLabel is the custom-capture path; it is not the fit when a public baseline already matches your robot or when you want a fully managed enterprise program.

Robotics datasets

Robot training data marketplace

A robot training data marketplace coordinates demonstrations, trajectories, video, robot state, action streams, and evaluation sets across candidate suppliers. Truelabel converts buyer requirements into supplier-facing specs, routes them to candidate capture partners for review, and requires sample review before scale. The sourcing decision has four paths — open benchmark, internal lab, managed data vendor, or a marketplace that routes your spec to reviewed capture partners. Which is "best" depends on whether your bottleneck is a baseline, rights, niche capture, or speed. TrueLabel is the marketplace path; it is not the right path when a public benchmark already fits or when you want a fully managed enterprise program with no supplier selection.

Vision-language-action models

VLA training data

VLA training data is the set of synchronized triplets a vision-language-action policy learns from: a visual observation, a language instruction, and the executed action at each timestep. The mistake teams make when they buy it is treating it as one homogeneous asset. It isn't. A workable VLA dataset is a mix — egocentric human video for affordance and dynamics pretraining, sensorized human demonstrations for action structure, robot or teleoperation data for embodiment-aligned actions, and a separate held-out set for evaluation. The reason the mix matters is evidence-backed: OpenVLA, a 7B model trained on 970,000 robot episodes, reports outperforming the 55B RT-2-X by 16.5% absolute in its own evaluation (OpenVLA) — source-specific evidence that a well-chosen data mix can outweigh raw model scale in that evaluation, not a universal law that diversity always beats parameter count. But its model card is equally clear that zero-shot use is bounded by the embodiments and domains in that mix (model card). TrueLabel is a physical AI data marketplace: you post a spec, vetted suppliers return samples. This page explains how to specify the mix for a VLA program, what each stream can and cannot supply, and how to scope a rights-cleared sample packet before you fund scale. For the full conceptual model, see the integrative guide on how VLAs, world models, and egocentric data fit together.

Fast validation

Eval data for robotics

Robotics eval data is a smaller dataset used to test model behavior, supplier quality, or task coverage before a larger training-data buy. truelabel's eval-request path lets buyers source a pre-scoped sample set with rights, consent, metadata, and acceptance criteria attached.

Buyer intake

Depth data for robotics

Depth data for robotics is per-pixel distance information — RGB-D frames, LiDAR point clouds, or learned monocular depth maps — that tells a policy how far every surface is, not just what it looks like. Robots use it for grasp planning and collision checking in manipulation, and for obstacle avoidance and SLAM in navigation. It comes in three shapes with very different accuracy, cost, and licensing: sensor RGB-D from active-stereo or structured-light cameras (Intel RealSense, Kinect v2, Zed), sparse LiDAR point clouds, and dense monocular depth predicted from ordinary RGB by models like Depth Anything V2. Most public RGB-D corpora are research-licensed, so commercial training usually needs custom capture with consent and commercial rights. Truelabel matches buyer depth specs to candidate capture suppliers reviewed against the buyer spec and delivers depth-channel data in RLDS, LeRobot, and MCAP with camera intrinsics/extrinsics, per-frame color-depth alignment, and rights-cleared provenance.

LeRobot dataset catalog

LeRobot datasets: browse the LeRobot-format catalog

LeRobot datasets are robot-learning datasets published in the LeRobotDataset format on the Hugging Face Hub: Parquet tables for states and actions, MP4 for camera frames, and JSON metadata for episode boundaries. Hundreds are public — BridgeData, DROID, RT-1/fractal, ALOHA sim, PushT — but licenses and embodiments vary, so match each set to your task before you train.