Fast validation
Eval data for robotics
Robotics eval data is a smaller dataset used to test model behavior, supplier quality, or task coverage before a larger training-data buy. truelabel's eval-request path lets buyers source a pre-scoped sample set with rights, consent, metadata, and acceptance criteria attached.
Quick facts
- Request type
- EVAL
- Scope
- Small fixed bundle for review
- Data
- Egocentric, teleop, manipulation, or custom modality
- Turnaround
- Short pilot before larger capture
- Acceptance
- Buyer reviews sample against checklist
Comparison
| Use case | Why eval first | Next step |
|---|---|---|
| New supplier | Validate quality before scale | Convert to OTS or net-new sourcing |
| New modality | Check format and QA assumptions | Refine specs |
| Model benchmark | Create a small held-out set | Request larger eval suite |
FAILURE & RECOVERY DATA
Failure and recovery demonstrations: a sourced atlas
Failure and recovery data is demonstrations of a robot getting a task wrong and correcting it. This atlas maps that emerging data class: every row ties a cited dataset or paper to failure type, task, embodiment, onset, root cause, recovery outcome, real or sim, modalities, and rights, with license and consent kept separate.
The eight-record sample is directional, not a complete or representative census. It covers the robot failure dataset, robot failure detection dataset, robot recovery data, failure-correction trajectories, and recovery demonstration dataset intents while keeping every unverified field explicit.
Exports: JSON · CSV. For comparison prose rather than atlas records, see common failure modes in ego/exocentric data.
| Dataset and source | Failure × task × embodiment | Labels and onset | Cause and recovery outcome | Environment, modalities, format | License and consent | Evidence status |
|---|---|---|---|---|---|---|
| FailSafe arXiv 2510.01642 · paper · Abstract and method sections | Generated execution failures Robot manipulation; exact task set is source-defined · Unknown — not verified in the registered primary source | Failure reasoning and executable recovery; success/suboptimal labels unverified Onset: Generated failure point | Source-generated failure reason Outcome: Executable recovery data reported | sim Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | License: Unknown — not verified in the registered primary source Consent: Not applicable to generated simulation unless a real-data input subset is used; verify inputs | needs-review · medium Claim: lane02-failsafe-generation · Entity: failsafe · Field: failure/recovery coverage · Unit: categorical · paper-arxiv-2510-01642 · 2026-07-22 · unverified — source retrieval hash not recorded · Source-reported, not independently validated; unknown fields require direct publisher review. |
| RePO-VLA / FRBench arXiv v1 2605.09410 · paper · Abstract, benchmark, and error-injection sections | Structured injected errors Recovery benchmark tasks defined by FRBench · Unknown — not verified in the registered primary source | Failure and recovery benchmark labels; suboptimal/success granularity is source-defined Onset: Structured error injection point | Injected benchmark error category Outcome: Benchmark recovery outcome | sim Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | License: Unknown — not verified in the registered primary source Consent: Not reported in the paper landing page; verify any real-data inputs | needs-review · medium Claim: lane02-frbench-error-injection · Entity: repo-vla-frbench · Field: failure/recovery coverage · Unit: categorical · paper-arxiv-2605-09410 · 2026-07-22 · unverified — source retrieval hash not recorded · Source-reported, not independently validated; unknown fields require direct publisher review. |
| ARMOR ICLR 2026 publisher PDF · paper · Abstract and failure detection/reasoning method | Robot execution failure detection Robotic task executions reported by the paper · Unknown — not verified in the registered primary source | Failure detection and reasoning; executable recovery labels not established Onset: Detected failure point | Vision-language failure reasoning Outcome: No executable recovery trajectory verified | unknown Vision and language reasoning · Unknown — not verified in the registered primary source | License: Unknown — not verified in the registered primary source Consent: Unknown — not verified in the registered primary source | needs-review · medium Claim: lane02-armor-detection-reasoning · Entity: armor · Field: failure/recovery coverage · Unit: categorical · paper-amazon-science-armor-iclr-2026 · 2026-07-22 · unverified — source retrieval hash not recorded · Source-reported, not independently validated; unknown fields require direct publisher review. |
| ProbeAct arXiv 2606.09740 · paper · Abstract and recovery method | Online execution failure Robot tasks evaluated by the paper · Unknown — not verified in the registered primary source | Failure and recovery outcomes; success/suboptimal label schema unverified Onset: Unknown — not verified in the registered primary source | Unknown — not verified in the registered primary source Outcome: Training-free recovery reported | unknown Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | License: Unknown — not verified in the registered primary source Consent: Unknown — not verified in the registered primary source | needs-review · medium Claim: lane02-probeact-recovery · Entity: probeact · Field: failure/recovery coverage · Unit: categorical · paper-arxiv-2606-09740 · 2026-07-22 · unverified — source retrieval hash not recorded · Source-reported, not independently validated; unknown fields require direct publisher review. |
| Oopsie Project page checked 2026-07-22 · project · Project dataset description | Robot failure examples Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | Failure examples; success/suboptimal/recovery coverage unverified Onset: Unknown — not verified in the registered primary source | Unknown — not verified in the registered primary source Outcome: Unknown — not verified in the registered primary source | unknown Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | License: Unknown — not verified in the registered primary source Consent: Unknown — not verified in the registered primary source | needs-review · medium Claim: lane02-oopsie-failure-data · Entity: oopsie · Field: failure/recovery coverage · Unit: categorical · project-oopsie-data · 2026-07-22 · unverified — source retrieval hash not recorded · Source-reported, not independently validated; unknown fields require direct publisher review. |
| BotFails Hugging Face card checked 2026-07-22 · dataset · Dataset card | Robot failure records Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | Failure labels source-reported; recovery and suboptimal coverage unverified Onset: Unknown — not verified in the registered primary source | Unknown — not verified in the registered primary source Outcome: Unknown — not verified in the registered primary source | unknown Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | License: Unknown — not verified in the registered primary source Consent: Unknown — not verified in the registered primary source | needs-review · medium Claim: lane02-botfails-card · Entity: botfails · Field: failure/recovery coverage · Unit: categorical · dataset-huggingface-botfails · 2026-07-22 · unverified — source retrieval hash not recorded · Source-reported, not independently validated; unknown fields require direct publisher review. |
| AHA Project page checked 2026-07-22 · project · Project overview | Failure analysis examples Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | Failure analysis source-reported; executable recovery unverified Onset: Unknown — not verified in the registered primary source | Analysis labels reported by publisher Outcome: Unknown — not verified in the registered primary source | unknown Vision-language annotations; exact streams unverified · Unknown — not verified in the registered primary source | License: Unknown — not verified in the registered primary source Consent: Unknown — not verified in the registered primary source | needs-review · medium Claim: lane02-aha-analysis · Entity: aha · Field: failure/recovery coverage · Unit: categorical · project-aha-vlm · 2026-07-22 · unverified — source retrieval hash not recorded · Source-reported, not independently validated; unknown fields require direct publisher review. |
| RoboFAC Hugging Face card checked 2026-07-22 · dataset · Dataset card | Robot failure and correction records Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | Failure/correction coverage source-reported; onset and recovery execution unverified Onset: Unknown — not verified in the registered primary source | Unknown — not verified in the registered primary source Outcome: Unknown — not verified in the registered primary source | unknown Unknown — not verified in the registered primary source · Unknown — not verified in the registered primary source | License: Unknown — not verified in the registered primary source Consent: Unknown — not verified in the registered primary source | needs-review · medium Claim: lane02-robofac-card · Entity: robofac · Field: failure/recovery coverage · Unit: categorical · dataset-huggingface-robofac · 2026-07-22 · unverified — source retrieval hash not recorded · Source-reported, not independently validated; unknown fields require direct publisher review. |
Next steps: compare public corpora, check dataset fit, or request held-out failure and recovery data. Sample packets come before scale, with rights, consent, and per-trajectory provenance attached.
When to start with eval data
Start with eval data when the buyer needs evidence quickly: a sample of supplier quality, a held-out benchmark, or a narrow slice of a larger environment before committing to a full capture program. Real-world robot datasets such as DROID show how much scene and task diversity can matter before scale [1], while Open X-Embodiment shows why cross-robot format coverage should be checked early [2]. CALVIN-style zero-shot evaluation also makes held-out language, environment, and object conditions explicit before a buyer treats a sample as production-ready [3].
[4]"Binary success is too coarse for long-horizon control, we therefore use a stage-wise scoring scheme."
That is the signal an eval request should create: not just pass/fail, but enough scoring detail to decide whether to refine the spec, change supplier, or scale capture.
What makes eval data useful
Useful eval sets are small but specific. They should preserve the metadata, rights, consent, and QA standards expected from a larger dataset so the buyer can trust the signal before scaling. Stress cases matter: THE COLOSSEUM shows that environmental perturbations can cut manipulation success rates sharply [5], and ManipArena argues real-world execution exposes perception noise, contact dynamics, hardware constraints, and latency that simulator-only eval misses [6]. Delivery should preserve synchronized logs and schemas, with formats like expert-led model evaluation workflows or MCAP-style timestamped multimodal data keeping review artifacts attached to the eval result [7].
Eval data vs training data
Training data teaches a model; eval data measures whether it generalizes. Eval sets should be held out, deduplicated from training, and designed around task success, failure modes, environment splits, and scoring rules rather than simply being more examples from the same distribution. Teams buying VLA training data, teleoperation traces, or robot demonstrations should define the eval split before suppliers scale collection.
| Dimension | Training data | Eval data |
|---|---|---|
| Purpose | Fit model behavior | Measure generalization and supplier quality |
| Overlap | May include many similar examples | Must avoid leakage and duplicates |
| Labels | Actions, states, demonstrations | Success/failure, scoring rubric, edge cases |
| Decision | Improve model | Accept, reject, or regress a model/data supplier |
Robotics eval dataset taxonomy
Useful eval bundles include supplier QA samples, held-out deployment scenes, regression suites, perturbation/OOD sets, safety or edge-case reviews, format-validation sets, VLA instruction-following evals, and navigation route evals. Each needs required fields, labels, metrics, and rejection reasons. A format-validation eval should load the sample in the LeRobot format or target schema before any broader buy is approved.
Leakage-resistant eval design
Hold out environments, objects or SKUs, contributors, time periods, lighting conditions, route IDs, instruction paraphrases, and failure cases. Ask suppliers to attest that eval examples were excluded from training and provide enough metadata to audit overlap. For marketplace requests, include the leakage audit in the robot training data marketplace acceptance rubric and keep rights review linked through egocentric data licensing when people or workplaces appear.
Eval-set acceptance checklist
A quotable robotics eval bundle should include held-out split definition, task instructions, embodiment, environment/object split keys, metric or rubric, failure taxonomy, leakage attestation, timestamps, action/state schema when applicable, provenance files, and a small load test. Reject eval samples that duplicate training scenes, omit failure cases, hide labels, or cannot be scored by an independent reviewer.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
External references and source context
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
DROID reports 350 hours of robot manipulation data across 86 tasks.
arXiv ↩ - Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Open X-Embodiment provides standardized robotic learning data across many robots, skills, and tasks.
arXiv ↩ - CALVIN paper
CALVIN evaluates agents zero-shot on novel instructions, environments, and objects.
arXiv ↩ - LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks
LongBench uses stage-wise scoring because binary success is too coarse for long-horizon control.
arXiv ↩ - THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
THE COLOSSEUM evaluates manipulation models across 14 environmental perturbation axes.
arXiv ↩ - ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
ManipArena targets real-world evaluation gaps caused by perception noise, contact dynamics, hardware, and latency.
arXiv ↩ - MCAP file format
MCAP stores multiple channels of timestamped multimodal log data for robotics applications.
mcap.dev ↩ - truelabel egocentric data glossary
Internal contextual link to the egocentric data definition.
truelabel.ai - truelabel sourcing brief intake
Internal contextual link to Truelabel's sourcing intake workflow.
truelabel.ai - truelabel egocentric warehouse video sourcing spec
Internal contextual link to warehouse egocentric video sourcing.
truelabel.ai - truelabel egocentric kitchen video sourcing spec
Internal contextual link to kitchen egocentric video sourcing.
truelabel.ai - truelabel industrial egocentric video sourcing spec
Internal contextual link to industrial egocentric video sourcing.
truelabel.ai - truelabel warehouse robotics data sourcing
Internal contextual link to warehouse robotics data sourcing.
truelabel.ai - truelabel kitchen manipulation data sourcing
Internal contextual link to kitchen manipulation data sourcing.
truelabel.ai - truelabel LeRobot dataset alternative comparison
Internal contextual link to the LeRobot dataset alternative comparison.
truelabel.ai - truelabel eval data for robotics hub
Internal contextual link to robotics eval data sourcing.
truelabel.ai - truelabel hand-object interaction data page
Internal contextual link to hand-object interaction training data requirements.
truelabel.ai - truelabel egocentric video datasets hub
Internal contextual link to the egocentric video datasets hub.
truelabel.ai
FAQ
What is an eval data sourcing request?
An eval data request is a smaller request for data that helps a buyer evaluate model behavior, supplier quality, or task coverage before funding a larger training-data program.
Is eval data exclusive?
Eval requests can be configured as exclusive or non-exclusive depending on the buyer's requirements and the supplier's rights model. The request should state this clearly before samples are reviewed.
What should an eval bundle include?
An eval bundle should include the data files, required metadata, consent artifacts where applicable, sample-level notes, and a clear acceptance checklist tied to the buyer's model or QA question.
Can eval data become a larger collection program?
Yes. A successful eval request is often the fastest way to refine the spec and then fund a larger off-the-shelf or net-new collection program.
How is robotics eval data different from training data?
Training data teaches the model; eval data measures generalization. Eval sets should be excluded from training and deduplicated from training sources to avoid leakage.
What should be included in a robotics eval bundle?
Include task definitions, robot embodiment, sensor streams, timestamps, action/state schema when applicable, success labels, failure reasons, environment/object splits, consent/provenance artifacts, scoring rubric, and a loadable manifest.
Looking for eval data for robotics?
Specify modality, task, environment, requested rights posture, and delivery format. Truelabel routes the request to candidate capture partners and helps scope consent/provenance artifacts and commercial licensing requirements for buyer review before delivery.
Request eval data