truelabelRequest dataEarnRequest

Vision-Language-Action Model

Pi-0.5 Training Data Requirements & Multi-Embodiment Dataset Specifications

Pi-0.5 is Physical Intelligence's 3-billion-parameter vision-language-action model released February 2025, trained on over 10,000 robot demonstration hours across 24 embodiments. It uses FAST action tokenization to convert continuous 7–24 DoF trajectories into discrete tokens, accepts multi-view RGB at 50 Hz plus proprioceptive state, and predicts 50-step action chunks per forward pass. Training requires hardware-synchronized camera streams, smooth teleoperation trajectories that yield clean FAST codebooks, and natural-language task instructions, often with chain-of-thought reasoning annotations for long-horizon tasks.

Updated 2026-07-1411 min read
By Truelabel Team
Reviewed by Truelabel Team ·
pi-0.5 training data

Quick facts

Topic
PI 0 5
Audience
Procurement leads, ML ops, robotics engineers
Deliverable
Buyer-facing reference + procurement guidance

What Is Pi-0.5 and Why It Matters for Physical AI

Pi-0.5 is a 3-billion-parameter vision-language-action (VLA) model from Physical Intelligence, the successor to pi-zero. The design bet is tokenization. Instead of a flow-matching or diffusion action head, pi-0.5 uses FAST (Factorized Action Sequence Tokenization) to turn continuous robot actions into discrete codebook tokens[1]. Actions then predict the way text does, one token at a time, so a single next-token objective covers both internet vision-language data and robot demonstrations without a separate action module.

That is why the interesting constraint sits on data, not architecture. Physical Intelligence reports over 10,000 hours across 24 embodiments[1], and reproducing that on your own task means smooth, hardware-synchronized demonstrations that survive FAST tokenization. Truelabel's physical-AI data marketplace matches that spec with calibrated teleoperation rigs, frame-synced multi-view RGB at 50 Hz, and reasoning-trace annotation for long-horizon tasks.

Inputs, Outputs, and the Action Codebook

Pi-0.5 conditions on three inputs: multi-view RGB, proprioceptive state, and a language instruction. Up to three cameras feed it, a third-person view plus optional left and right wrist cameras, each frame captured at 50 Hz and resized to a standard resolution (typically 224×224) for the PaLI-3 ViT encoder[1]. Proprioception is joint positions, velocities, and binary gripper state, projected into the token stream alongside the vision tokens.

The action space is continuous and embodiment-specific, from 7 DoF on a single arm to 24 DoF on a bimanual mobile base. Pi-0.5 never regresses those numbers directly. FAST maps trajectories into a discrete codebook via vector quantization, and at inference the transformer autoregressively emits 50 action tokens per forward pass, decoded back into a 50-step control chunk at 50 Hz[1]. Chunked prediction is what keeps compounding error low without collapsing to single-step reactivity.

Architecturally the encoder is a ViT-G/14 that turns each 224×224 frame into 256 visual tokens; three views plus language and action tokens give a context on the order of 768 visual + 50 language + 50 action tokens per step, run through a decoder-only backbone with a unified causal mask[1]. One consequence matters for buyers: the codebook is shared across all 24 training embodiments, so a token index maps to different joint angles on a Franka Panda than on a UR5. Fine-tune on one embodiment or pass an embodiment ID; Truelabel's LeRobot-compatible format ships per-episode embodiment metadata so either path works without post-processing.

ChannelShape and rateNotes
RGB viewsup to 3 cameras, 224×224, 50 HzThird-person plus optional L/R wrist; PaLI-3 ViT-G/14, 256 tokens per frame
Proprioceptionjoint pos + vel + gripper, 50 HzAbout 15 dimensions on a 7-DoF arm
Languageinstruction stringTask text, optional chain-of-thought
Action7 to 24 DoF, 50-token chunkFAST codebook, decoded to 50 steps at 50 Hz
Pi-0.5 observation and action interface

Why FAST Punishes Jerky Data

FAST is a learned VQ-VAE. During pre-training it fits a codebook of 256 to 1024 motion prototypes (reach, grasp, retract, place), sized per embodiment[1]. That learned codebook is also the failure mode. Jerky or mode-switching teleoperation produces prototypes that do not generalize, and the policy then fails on out-of-distribution workspace layouts. Velocity spikes and high-frequency jitter degrade token quality directly[2].

Camera timing is the second silent killer. If views drift apart or frames drop, the 50 Hz action-observation correspondence breaks and supervised training learns the wrong pairing. Truelabel delivers frame-synced multi-view RGB with intrinsic and extrinsic calibration and microsecond timestamps, plus QA evidence and sample packets, so a buyer can audit smoothness and alignment before ordering at scale.

Training-Data Scale vs Public Benchmarks

Pi-0.5's 10,000-plus hours across 24 embodiments is the largest proprietary robot set disclosed to date[1]. For scale intuition, Open X-Embodiment aggregates roughly 1 million episodes (about 2,000 hours) from 22 robots, and DROID adds 76,000 episodes (about 350 hours) on 8 Franka Panda cells. Pi-0.5 therefore trained on roughly 5 times the hours of the largest public aggregate, skewed toward household manipulation and light assembly.

The distribution matters more than the total. Coverage spans single-arm (Franka, UR5, Kinova Gen3), bimanual (ALOHA, dual Franka), mobile (Fetch, TIAGo), and dexterous hands (Allegro, Shadow)[1]. Single-robot datasets often fail to transfer even to kinematically similar arms, because gripper geometry, joint limits, and control latency differ. For a custom task Physical Intelligence recommends 500 to 2,000 demonstrations for reliable closed-loop performance[1]. Truelabel's marketplace supplies pre-collected task libraries or custom collection on your embodiment, including embodiment-transfer pairs (the same task on two robots) for studying how codebook statistics shift.

CorpusScaleEmbodimentsLicense
Pi-0.5 (Physical Intelligence)10,000+ hours24Proprietary
Open X-Embodiment~1M episodes (~2,000 hours)22 robotsMixed, per constituent
DROID76,000 episodes (~350 hours)8 Franka cellsCC BY 4.0
Proprietary pi-0.5 data vs public robot-demonstration corpora

How Pi-0.5 Differs from RT-2 and OpenVLA

RT-2 pairs a PaLI-X encoder with a discretized-token action head—it emits actions as text tokens that de-tokenize into discrete action bins—and needs 130,000-plus demonstrations for reliable behavior[3]. OpenVLA is a 7B open-source model trained on the roughly 1-million-episode Open X-Embodiment set[4]. Pi-0.5's differentiator is the learned action codebook: because actions are tokens, one next-token objective jointly pre-trains on web image-text corpora and robot demos, and multi-embodiment training uses a single shared codebook with conditional decoding rather than a regression head per action space.

The trade-off is coverage. If the training distribution lacks smooth trajectories for a motion, such as high-speed throwing or contact-rich insertion, the VQ-VAE cannot reconstruct that primitive and execution fails. So pi-0.5 is strong out of the box on household manipulation and weak on industrial or outdoor tasks that its 10,000-hour set underrepresents[1]. Closing that gap is a data problem: Truelabel takes domain-specific bounties (workspace, success criteria, embodiment) and returns demos formatted for immediate fine-tuning.

Collecting Pi-0.5-Ready Data

A pi-0.5-ready episode is time-aligned RGB, proprioception, action, and language in one container, and the requirements shift by embodiment. The canonical camera setup is a static third-person view 1 to 1.5 m from the workspace plus two wrist cameras, all frame-synced to within 5 ms through a hardware trigger or shared clock, or fused features smear[2]. Single arms (Franka, UR5, Kinova) give 7 DoF and mature ROS integration; Truelabel logs 50 Hz state over the Franka FR3 Control Interface with three RealSense cameras. Bimanual ALOHA doubles the space to 14 DoF and needs backdrivable, gravity-compensated leaders so the operator guides both follower arms without fighting the motors. Mobile bases (Fetch, TIAGo, Spot with arm) reach 10 to 24 DoF and add navigate-grasp-retreat sequencing over ROS navigation on pre-mapped scenes, at several times the per-hour cost of fixed-base capture.

  1. 01

    Lock camera sync

    One hardware trigger or PTP clock across the third-person and both wrist views, verified against dropped frames when USB bandwidth saturates.

  2. 02

    Ship calibration per episode

    Intrinsic and extrinsic matrices for every camera, so point-cloud reconstruction and multi-view consistency losses stay possible downstream.

  3. 03

    Log state and action together

    Joint positions, velocities, gripper, and end-effector pose on the same timestamp array as the frames.

  4. 04

    Keep trajectories smooth

    No velocity spikes or mode switches; smoothness is what makes FAST tokens generalize across new layouts.

  5. 05

    Attach language and success labels

    Task instruction, optional chain-of-thought, and a human-verified success flag per episode.

Delivery Formats: HDF5, RLDS, and LeRobot

Truelabel ships pi-0.5-compatible episodes as HDF5 or RLDS, convertible to LeRobot. One HDF5 file is one episode: `/observations/images/{primary,wrist_left,wrist_right}` (T×224×224×3 uint8), `/observations/proprioceptive_state` (T×15 float32 for a 7-DoF arm), `/actions` (T×7 float32), `/language_instruction`, and `/metadata` (embodiment, task, collector, timestamp)[5]. For the openpi framework, RLDS wraps each episode as a `tf.data.Dataset` with standard keys (observation, action, reward, discount, language_instruction) that drop into existing pipelines.

LeRobot conversion matters when you mix Truelabel data with public sets like ALOHA or DROID. Episodes become Parquet with frame-level timestamps in a private Hugging Face repo, loadable via `lerobot.datasets.load_dataset(...)` and trainable with LeRobot's Diffusion Policy or ACT implementations before you switch to FAST tokenization. Delivery targets RLDS, LeRobot, MCAP, or a custom schema on request.

Marketplace, Provenance, and Licensing

Buyers post a spec (embodiment, workspace, success criteria, annotation schema) and matched suppliers return sample packets before any large order. Three shapes cover most pi-0.5 work: catalog task libraries for immediate HDF5 or RLDS download, custom collection on your embodiment and workspace, and hybrid runs that reuse catalog primitives like pick-and-place while capturing the domain-specific parts like PCB assembly. Every demonstration carries per-trajectory provenance records with collector ID, timestamp, embodiment, and calibration reference, which is the paper trail teams need for AI Act Article 10 or NIST AI RMF data-governance documentation.

Licensing is where public data trips procurement. Pi-0.5's own weights are proprietary and cannot be redistributed. Truelabel data ships under CC BY 4.0 or custom commercial terms that grant perpetual, worldwide train-and-deploy rights with no revenue share and no non-commercial clause[6]. Mixing in public sets needs an audit: Open X-Embodiment bundles 22 datasets under CC BY 4.0, CC BY-NC 4.0, MIT, and Apache 2.0, so one non-commercial constituent can taint a release[7], and DROID is CC BY 4.0 but its README adds an informal citation request that some legal teams treat as a moral obligation[2]. For restricted environments, Truelabel also runs on-premises collection under buyer supervision.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

External references and source context

  1. General agents need world models

    Pi-0.5 model architecture, 10,000-hour training set, 24-embodiment coverage, and FAST tokenization design

    arXiv ↩
  2. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    DROID dataset scale (76,000 episodes, 350 hours), 8 Franka Panda setups, and camera synchronization requirements

    arXiv ↩
  3. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    RT-2 model comparison, 130,000-demonstration training scale, and encoder-decoder VLA design

    arXiv ↩
  4. OpenVLA: An Open-Source Vision-Language-Action Model

    OpenVLA 7B-parameter model, Open X-Embodiment training, and open-source VLA benchmarks

    arXiv ↩
  5. RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning

    RLDS dataset format, TensorFlow Datasets integration, and episode schema standardization

    arXiv ↩
  6. Attribution 4.0 International deed

    Creative Commons Attribution 4.0 license terms and commercial use permissions

    Creative Commons ↩
  7. Open X-Embodiment: Robotic Learning Datasets and RT-X Models

    Open X-Embodiment dataset aggregation, 970,000 episodes, 22 robots, and multi-embodiment training

    arXiv ↩
  8. BridgeData V2: A Dataset for Robot Learning at Scale

    BridgeData V2 dataset scale, task horizon comparison, and single-arm manipulation benchmarks

    arXiv
  9. scale.com physical ai

    Scale AI physical-AI data engine and proprietary dataset infrastructure

    scale.com
  10. docs.labelbox.com overview

    Labelbox annotation platform and multi-sensor data labeling workflows

    docs.labelbox.com
  11. encord.com active

    Encord Active data curation and quality control pipelines

    encord.com
  12. RoboNet: Large-Scale Multi-Robot Learning

    RoboNet multi-robot dataset scale and embodiment diversity

    arXiv
  13. Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

    EPIC-KITCHENS egocentric video dataset and long-horizon task annotation

    arXiv
  14. Datasheets for Datasets

    Datasheets for Datasets framework and dataset documentation best practices

    arXiv
  15. Model Cards for Model Reporting

    Model Cards for Model Reporting and transparency requirements

    arXiv
  16. Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World

    Domain randomization for sim-to-real transfer and synthetic data augmentation

    arXiv
  17. CALVIN paper

    CALVIN benchmark for long-horizon language-conditioned manipulation

    arXiv
  18. RT-1: Robotics Transformer for Real-World Control at Scale

    RT-1 Robotics Transformer architecture and Google robot fleet training

    arXiv
  19. RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

    RoboCat self-improving generalist agent and multi-task learning

    arXiv
  20. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

    SayCan language grounding in robotic affordances and task planning

    arXiv

FAQ

What is the minimum dataset size required to fine-tune pi-0.5 on a custom task?

Physical Intelligence recommends 500–2,000 demonstrations per task (25–100 hours at 50 Hz) for reliable closed-loop performance. Tasks with high variability (e.g., deformable manipulation, liquid pouring) require the upper end of this range, while constrained pick-and-place tasks land at the lower end. Truelabel's custom collection service delivers demonstrations with QA evidence for buyer review and per-trajectory provenance metadata.

Can I use pi-0.5 with a robot embodiment not included in Physical Intelligence's 24-embodiment training set?

Yes, but you will need to fine-tune on demonstrations from your target embodiment. Pi-0.5's FAST codebook is learned from Physical Intelligence's proprietary 24-embodiment mixture, so zero-shot transfer to a new kinematic chain (e.g., a 6-DoF UR3 or a 12-DoF humanoid arm) will likely fail due to action-space mismatch. Truelabel offers embodiment-transfer datasets: we collect 200–500 demonstrations on your robot performing the same tasks as a Franka Panda or ALOHA, enabling you to learn a new FAST codebook while preserving pi-0.5's visual and language representations.

How does FAST tokenization compare to diffusion-based action prediction (e.g., Diffusion Policy)?

FAST discretizes continuous actions into a learned codebook and predicts tokens autoregressively, while Diffusion Policy iteratively denoises a Gaussian action distribution over 10–100 steps. FAST is faster at inference (single forward pass vs. 10–100 denoising steps) and scales better to web-scale pre-training (next-token prediction is the same objective as language modeling), but it requires smooth training trajectories to learn a high-fidelity codebook. Diffusion Policy is more robust to noisy demonstrations but cannot leverage internet vision-language data without architectural modifications. For buyers with clean teleoperation data and access to large-scale compute, FAST offers better sample efficiency; for buyers with lower-quality data or limited compute, Diffusion Policy may be more practical.

What camera hardware does truelabel use for 50 Hz multi-view RGB collection?

Truelabel collectors use Intel RealSense D435 or D455 cameras for wrist views (compact, USB 3.0, built-in IMU) and Basler ace or FLIR Blackfly cameras for third-person views (higher resolution, GigE interface, hardware trigger support). All cameras are synchronized via a shared hardware trigger or PTP (Precision Time Protocol) clock, ensuring frame alignment within 5 ms. We deliver intrinsic and extrinsic calibration matrices for every episode, enabling downstream 3D reconstruction or multi-view consistency losses.

What does truelabel deliver for FAST codebook training on common embodiments?

Truelabel delivers smooth, 50 Hz demonstration datasets for embodiments like Franka Panda (7 DoF), UR5 (6 DoF), and ALOHA (14 DoF), formatted so buyers can train pi-0.5's FAST codebook and action decoder on their target embodiment. Each dataset ships with per-trajectory provenance metadata and QA evidence for buyer review, and sample packets let you confirm that your target task's action distribution is well covered before committing to a full order.

How does truelabel handle task success labeling for ambiguous or multi-stage tasks?

Human annotators watch a sped-up replay of each episode and mark binary success (task completed) or failure (task not completed). For multi-stage tasks (e.g., "make a sandwich"), we also annotate sub-goal completion (bread placed, condiments spread, sandwich assembled) to enable hierarchical policy training. Annotators are trained on task-specific rubrics with photo examples of success/failure states, and we enforce inter-rater reliability by having 10% of episodes labeled by two independent annotators. If their labels disagree, a senior annotator reviews the episode and makes a final determination.

Looking for pi-0.5 training data?

Specify modality, task, environment, requested rights posture, and delivery format. Truelabel routes the request to candidate capture partners and helps scope consent/provenance artifacts and commercial licensing requirements for buyer review before delivery.

Submit Pi-0.5 Dataset Bounty