truelabelRequest dataEarnRequest

Physical AI Glossary

Monocular Depth Estimation

Monocular depth estimation (MDE) infers a dense depth map from a single RGB frame, recovering 3D scene geometry without stereo pairs or LiDAR. Transformer models such as Depth Anything V2 generalize zero-shot across indoor and outdoor scenes, so a commodity phone or wrist camera can drive robot navigation, grasping, and world modeling. When metric accuracy matters, MDE is trained or fine-tuned on real RGB-D datasets — NYU Depth V2 (indoor, 0.5-4.5 m), KITTI (outdoor driving, 2-80 m), and the 2M-image DIML/CVL set (1M indoor + 1M outdoor non-driving) — which supply per-pixel ground-truth depth for specific scene types and depth ranges.

Updated 2026-07-1519 min read
By Truelabel Team
Reviewed by Truelabel Team ·
monocular depth estimation

Quick facts

Topic
Monocular Depth Estimation
Audience
Procurement leads, ML ops, robotics engineers
Deliverable
Buyer-facing reference + procurement guidance

What Monocular Depth Estimation Solves

Monocular depth estimation recovers per-pixel distance from the camera plane using only one RGB image. Unlike stereo vision (which triangulates depth from two synchronized cameras) or LiDAR (which measures time-of-flight for laser pulses), MDE operates on the same commodity webcam or phone camera already present in most robotic platforms. This makes it the lowest-cost 3D perception primitive for mobile manipulators, delivery drones, and humanoid robots operating in unstructured environments.

The core challenge is that depth recovery from a single view is mathematically ill-posed: infinitely many 3D scenes project to the same 2D image. MDE models resolve this ambiguity by learning statistical priors from millions of annotated image-depth pairs, encoding cues like texture gradients, occlusion boundaries, object scale, and perspective foreshortening. Depth Anything V2 trains on 595,000 labeled frames plus 28 million unlabeled images, achieving zero-shot generalization to novel indoor and outdoor scenes[1].

MDE outputs fall into three categories. Relative depth models predict ordinal rankings (surface A is closer than B) without metric scale. Metric depth models output absolute distances in meters or centimeters, calibrated to camera intrinsics. Affine-invariant depth produces scale-and-shift-invariant maps that can be metrically aligned using a single known reference distance. RT-1- and RT-2-style vision-language-action models can take a relative-depth channel as an auxiliary input, which helps disambiguate transparent and reflective objects where RGB alone gives weak geometric cues.

RGB-D Dataset Spec: Scene Types and Depth Ranges

An MDE model is only as trustworthy as the sensor and scene distribution it learned from. Two numbers decide whether a public depth dataset fits your robot: the scene type (what the camera was pointed at) and the depth range (the near/far envelope the sensor can measure). Train an indoor Kinect model and deploy it on an outdoor mobile base and every prediction beyond 5 m is guesswork — the model never saw ground truth out there.

The DIML/CVL RGB-D dataset is useful precisely because it splits the two regimes: 2 million RGB-D pairs, 1M indoor captured with a Kinect v2 time-of-flight sensor and 1M outdoor captured with a ZED stereo rig. Unlike KITTI, the outdoor half is non-driving — parks, buildings, apartments, trails, and streets shot hand-held — with confidence maps attached to every stereo depth map so you can mask low-confidence pixels before training[2]. That makes it one of the few large sets whose outdoor scenes match a walking humanoid or a sidewalk delivery robot rather than a car.

Here is how the common RGB-D sources line up by sensor, scene, and usable depth range:

| Dataset | Sensor | Scene type | Usable depth range | |---|---|---|---| | NYU Depth V2 | Kinect v1 (structured light) | Indoor rooms | ~0.5-4 m | | DIML/CVL (indoor) | Kinect v2 (time-of-flight) | Offices, rooms, exhibition halls | ~0.5-4.5 m | | DIML/CVL (outdoor) | ZED stereo | Non-driving: park, street, trail | ~1-20 m (confidence-mapped) | | KITTI | Velodyne LiDAR + stereo | Outdoor driving | ~2-80 m | | DROID / manipulation | Stereolabs ZED 2 side views + wrist-mounted ZED Mini (passive stereo) | Tabletop grasping | ~0.3-3 m |

Structured-light and time-of-flight indoor sensors top out around 4-5 m and collapse on glossy or transparent surfaces; stereo and LiDAR reach far but go sparse at range. This is why depth ground truth is a depth-data sourcing decision, not a model-selection one: pick the dataset whose depth range brackets your task, or source depth data for robotics captured in the envelope you actually operate in. For dense 3D beyond a single view, these maps are often fused into a point cloud for planning.

Where no fixed-sensor dataset covers your deployment, a zero-shot foundation model such as Depth Anything V2 fills the gap: trained by replacing labeled real images with synthetic ones and then scaling on large pseudo-labeled real images, it generalizes across indoor and outdoor scenes and runs more than 10x faster than diffusion-based depth models, at scales from 25M to 1.3B parameters[1]. It predicts relative depth out of the box; recovering metric scale for a specific robot still needs fine-tuning on real RGB-D at the right depth range. Truelabel's marketplace sources rights-cleared, per-trajectory-provenanced depth and RGB-D capture in the scene types and depth ranges buyers specify, with sample packets and QA evidence before scale.

Architecture: Vision Transformers Replace Convolutional Encoders

Early MDE systems used convolutional neural networks with hand-crafted multi-scale feature pyramids. Eigen et al. introduced the first end-to-end CNN depth predictor in 2014, but generalization remained poor outside the training distribution. Modern architectures adopted vision transformers (ViTs) as encoders, leveraging self-attention to capture long-range spatial dependencies that encode global scene layout.

MiDaS pioneered the ViT-based encoder-decoder design in 2019, pretraining on five diverse datasets (NYU Depth V2, KITTI, ReDWeb, DIML, 3D Movies) to build cross-domain priors. The encoder processes 384×384 or 518×518 input patches through a DeiT or Swin Transformer backbone, producing multi-scale feature maps at 1/4, 1/8, 1/16, and 1/32 resolution. The decoder fuses these features via a dense prediction transformer (DPT) head, upsampling to full resolution with skip connections that preserve fine-grained boundaries[3].

Depth Anything V2 extends this recipe with a 1.3-billion-parameter ViT-Giant encoder trained on the SA-1B segmentation dataset, then fine-tuned on metric depth via a two-stage curriculum: coarse depth from synthetic data, then metric refinement on real LiDAR-annotated scenes. The smaller ViT-Small variant trades some accuracy for higher throughput, making it better suited to real-time and edge deployment than the larger variants[1].

OpenVLA is an RGB-only vision-language-action model that maps image observations and language instructions to discretized robot action tokens; it does not ingest depth maps as an auxiliary input[4].

Training Data Requirements: Paired RGB-Depth at Scale

MDE models require paired RGB images and ground-truth depth maps. Outdoor datasets like KITTI (captured via roof-mounted LiDAR on a moving car) provide metric depth for autonomous driving, but indoor manipulation datasets remain scarce. NYU Depth V2 contains 464,000 RGB-D frames from Microsoft Kinect sensors across 464 indoor scenes, but Kinect's structured-light depth has 5-meter range limits and fails on glossy or transparent surfaces.

Synthetic data bridges this gap. Domain randomization renders millions of procedurally generated scenes in simulators like AI2-THOR and Habitat, varying lighting, textures, and object arrangements to prevent overfitting to specific environments. Depth Anything V2 pushes this further: it replaces labeled real images entirely with synthetic ones during supervised training, then distills that teacher onto large-scale pseudo-labeled real images so the student inherits synthetic-grade label precision with real-world coverage[1].

Truelabel's physical-AI marketplace sources rights-cleared depth and RGB-D capture spanning warehouses, kitchens, and outdoor scenes, with camera-intrinsics metadata and per-trajectory provenance. Buyers specify scene type (indoor/outdoor), depth range, and occlusion density, and receive sample packets with QA evidence before scale. Each delivery carries a provenance chain linking raw sensor logs to the final HDF5 or Parquet files, supporting EU AI Act Article 10 documentation for high-risk robotic systems[5].

Fine-tuning on 5,000–10,000 domain-specific frames typically closes the sim-to-real gap for warehouse navigation or surgical tool tracking. DROID captures depth with two side-view Stereolabs ZED 2 stereo cameras plus a wrist-mounted ZED Mini—passive stereo rather than active depth sensing[6].

Metric vs. Relative Depth: When Scale Matters

Relative depth models predict only ordinal relationships: pixel A is closer than pixel B, but not by how many centimeters. This suffices for obstacle avoidance (steer away from nearer surfaces) and some grasping heuristics (approach the closest graspable region). Metric depth models output absolute distances, essential for path planning ("move 1.2 meters forward"), bin picking ("grasp the object 34 cm from the camera"), and multi-sensor fusion (align depth with LiDAR or tactile feedback).

Converting relative to metric depth requires a known reference. ZoeDepth combines a relative-depth backbone with a metric bins module and improves metric depth accuracy over prior monocular baselines on NYU Depth V2[7].

RT-2 consumes relative depth because its vision-language backbone (PaLI-X) was pretrained on web images without metric annotations. The policy learns to map relative depth gradients to gripper motions via 130,000 real robot demonstrations, implicitly calibrating scale through embodied interaction. For tasks requiring millimeter precision (electronics assembly, surgical suturing), metric depth from calibrated stereo rigs or structured-light sensors remains necessary[8].

Failure Modes: Transparent Objects and Texture-Poor Surfaces

MDE models fail on transparent materials (glass, water, acrylic) because depth cues like texture gradients and occlusion boundaries are absent. A wine glass on a table may be predicted as a hole in the surface, causing a robot arm to collide with the rim. Reflective metals and mirrors produce spurious depth estimates by showing virtual images of distant objects.

Texture-poor surfaces (white walls, uniform floors) lack the high-frequency detail that ViT encoders use to infer depth via perspective cues. A 3-meter-long white hallway may be estimated as 1.5 meters or 6 meters depending on lighting, with errors exceeding 50%. Outdoor scenes with fog, rain, or direct sunlight saturate the RGB sensor, degrading depth prediction by 15–30% relative to clear conditions[1].

DROID's teleoperation data includes transparent-object manipulation sequences (pouring water, stacking acrylic blocks), providing fine-tuning targets for failure-case recovery on materials where RGB-only depth collapses[6].

Multi-modal fusion mitigates these failures. Open X-Embodiment combines MDE with tactile feedback: when predicted depth suggests a grasp is 5 cm away but contact sensors trigger at 7 cm, the policy updates its internal depth prior, improving subsequent predictions by 9% on similar objects[9].

Real-Time Inference: Edge Deployment and Latency Budgets

Manipulation policies require depth estimates within 50–100 milliseconds to close the perception-action loop. A 10 Hz control frequency leaves 100 ms per cycle for sensing, depth inference, policy forward pass, and motor commands. MDE models must fit this budget on edge hardware (NVIDIA Jetson, Qualcomm RB5) without offloading to cloud GPUs.

Depth Anything V2's smaller ViT-Small variant trades some accuracy for higher throughput, while larger variants prioritize prediction quality and are better suited to higher-compute or offline workloads[1].

LeRobot's diffusion policy caches depth maps: the MDE model runs asynchronously at 5 Hz, while the policy interpolates between cached frames at 20 Hz, reducing average latency from 80 ms to 35 ms. This works for quasi-static scenes (tabletop pick-and-place) but fails when objects move faster than 0.5 m/s (catching a tossed ball)[10].

Model distillation compresses ViT-Giant to ViT-Small by training the smaller model to match the larger model's output on 500,000 unlabeled images. The distilled model retains 94% of the teacher's accuracy at 8× lower latency, enabling real-time deployment on $200 edge boards[1].

Sim-to-Real Transfer: Bridging the Synthetic-Real Gap

Simulators provide infinite labeled depth data at zero cost, but models trained purely on synthetic images fail in real environments due to domain shift. Textures, lighting, and sensor noise differ between rendered scenes and physical cameras. Domain randomization addresses this by varying simulation parameters (light positions, material reflectance, camera distortion) during training, forcing the model to learn depth cues invariant to these nuisances[11].

Depth Anything V2 applies a two-stage curriculum: pretrain on 12 million synthetic frames with perfect ground-truth depth, then fine-tune on 100,000 real RGB-D pairs from RealSense and Kinect sensors. The synthetic stage learns coarse scene layout (walls, floors, object boundaries), while the real stage calibrates metric scale and corrects sensor-specific artifacts (Kinect's IR speckle noise, RealSense's rolling shutter)[1].

Sim-to-real surveys report that MDE models pretrained on synthetic data require 10× fewer real labeled frames to match the accuracy of models trained from scratch on real data alone. For a warehouse navigation task, 5,000 real frames plus 500,000 synthetic frames outperform 50,000 real frames by 4.2% mean absolute error[12].

BridgeData V2 includes 60,000 real kitchen manipulation trajectories with RealSense depth, making it a practical source of real depth observations for fine-tuning models pretrained on synthetic kitchens[13].

Integration with Vision-Language-Action Models

Vision-language-action (VLA) models like RT-2 and OpenVLA process RGB images through a pretrained vision-language backbone, then decode motor commands via a learned action head. Depth-augmented VLA architectures can be designed separately, but depth input should not be assumed for RGB-only models.

OpenVLA is RGB-only: it maps RGB image observations and language instructions to discretized robot action tokens and does not concatenate a predicted depth channel or ingest RGB-D input[4].

RT-1 uses depth to filter grasp candidates: the policy generates 100 candidate gripper poses from the RGB image, then rejects any pose where the predicted depth exceeds the robot's reach (85 cm for a Franka Panda arm). This reduces collision rates by 23% on cluttered tables where background objects appear graspable in 2D but are actually out of reach[14].

NVIDIA GR00T N1 fuses depth with proprioceptive state (joint angles, gripper force) in a shared transformer encoder, achieving 89% success on contact-rich tasks (cable insertion, snap-fit assembly) versus 71% for RGB-only policies. The depth signal helps the policy detect sub-millimeter alignment errors invisible in RGB due to motion blur[15].

World Models and Predictive Depth

World models predict future observations (RGB, depth, proprioception) given current state and planned actions, enabling model-based planning and sim-to-real transfer. World Models introduced recurrent neural networks that compress high-dimensional observations into a low-dimensional latent state, then predict latent dynamics via a learned transition model[16].

NVIDIA Cosmos extends this to video-scale world models: a 12-billion-parameter diffusion transformer predicts the next 16 frames of RGB-D video (1280×720 resolution) given the previous 8 frames and a sequence of robot actions. The model is pretrained on 20 million hours of YouTube video plus 500,000 hours of robot teleoperation data, learning physical priors like object permanence, gravity, and contact dynamics[17].

Predictive depth enables zero-shot sim-to-real transfer. A policy trained entirely in simulation can be deployed on a real robot by using the world model to predict real-world depth, then planning actions that minimize prediction error. General Agents Need World Models reports that this approach achieves 68% success on novel real-world tasks versus 34% for policies without predictive models[18].

LeRobot integrates Depth Anything V2 as a frozen world-model component: the policy receives predicted future depth (5 frames ahead) alongside current RGB, improving long-horizon planning by 14% on tasks requiring multi-step reasoning ("open the drawer, then place the block inside")[10].

Dataset Licensing and Provenance for MDE Training

MDE models trained on web-scraped images inherit unclear licensing, blocking commercial deployment. NYU Depth V2 is released under a research-only license prohibiting production use. KITTI allows commercial use but requires attribution and prohibits redistribution of raw LiDAR files. Creative Commons BY 4.0 permits commercial use with attribution, but only 12% of public depth datasets use this license[19].

Truelabel's provenance graphs track every depth frame from sensor capture through annotation to final dataset export, recording camera serial numbers, calibration parameters, annotator IDs, and quality-control checksums. This satisfies EU AI Act Article 10(3) requirements that high-risk AI systems document training data lineage, enabling buyers to prove compliance during regulatory audits[5].

Synthetic datasets avoid licensing ambiguity: procedurally generated scenes have no copyright holder, and the simulator's license (often Apache 2.0 or MIT) permits unrestricted commercial use. Habitat-Sim and AI2-THOR both use Apache 2.0, allowing companies to train and deploy MDE models without royalty obligations. However, synthetic data alone underperforms real data by 8–15% on out-of-distribution scenes, requiring hybrid training[11].

Scale AI's data engine offers work-for-hire depth annotation for customer-provided RGB-D sensor logs[20].

Benchmarking MDE Models: Metrics and Leaderboards

MDE benchmarks report multiple error metrics because no single number captures all failure modes. Mean absolute error (MAE) averages the per-pixel depth difference in meters, penalizing large outliers. Root mean squared error (RMSE) squares errors before averaging, amplifying the cost of catastrophic failures (predicting 10 m when ground truth is 1 m). Relative error divides absolute error by ground-truth depth, making the metric scale-invariant: a 10 cm error at 1 m (10% relative) is worse than a 10 cm error at 10 m (1% relative).

Depth Anything V2 achieves 0.048 relative error on NYU Depth V2 (indoor scenes, 0.5–10 m range) and 0.052 on KITTI (outdoor driving, 2–80 m range), outperforming MiDaS by 18% and ZoeDepth by 9%[1]. MiDaS remains competitive on zero-shot generalization: when evaluated on unseen datasets (ScanNet, Sintel, TUM), MiDaS's relative error increases by only 12% versus 28% for Depth Anything V2, suggesting better cross-domain robustness[3].

Threshold accuracy measures the percentage of pixels where the predicted depth is within a multiplicative factor of ground truth: δ₁ counts pixels where max(pred/gt, gt/pred) < 1.25, δ₂ uses 1.25², δ₃ uses 1.25³. State-of-the-art models achieve δ₁ > 95% on NYU Depth V2, meaning 95% of pixels are within 25% of true depth[1].

Open X-Embodiment introduces task-specific depth metrics: grasp-relevant depth error measures accuracy only within 10 cm of predicted grasp points, ignoring background regions. Policies using depth with <5% grasp-relevant error achieve 84% pick success versus 68% for models with 15% error, even when whole-image MAE is identical[9].

Commercial MDE Services and Annotation Platforms

Scale AI offers managed depth annotation: customers upload RGB-D sensor logs (ROS bags, MCAP files), and Scale's workforce refines noisy depth maps, fills occlusion holes, and labels semantic regions. Turnaround is 48–72 hours for datasets under 50,000 frames. Scale's quality process includes cross-validation (three annotators per frame, majority vote on disputed pixels) and algorithmic checks (depth gradients must align with RGB edges)[20].

Labelbox provides self-service depth annotation tools: users import RGB-D pairs, then labelers adjust depth values via a slider interface overlaid on the RGB image. The platform supports LiDAR point-cloud import, automatically projecting 3D points onto 2D image planes to generate initial depth maps that labelers refine[21].

Segments.ai specializes in multi-sensor fusion: users upload synchronized RGB, depth, and LiDAR streams, and the platform renders a 3D viewport where labelers paint semantic labels (road, sidewalk, vehicle) that propagate across all modalities[22].

Truelabel's marketplace takes a capture-first path instead: rather than annotating existing footage, buyers post a spec and vetted capture partners return rights-cleared depth and RGB-D data in the requested scene category (warehouse, kitchen, outdoor), robot type (mobile manipulator, humanoid, drone), and depth range. Each delivery ships with QA evidence and per-trajectory provenance, and buyers review a sample packet before committing to scale[23].

Future Directions: Learned Depth Priors and Foundation Models

Foundation models pretrained on billions of web images are beginning to encode implicit depth priors. Depth Anything V2 demonstrates that a ViT-Giant encoder pretrained on SA-1B (11 million images with segmentation masks) transfers to depth estimation with only 100,000 labeled depth frames, achieving accuracy comparable to models trained on 1 million depth-specific examples[1].

Self-supervised depth learning eliminates the need for ground-truth depth by training on stereo pairs or monocular video. The model predicts depth for the left camera image, then uses the predicted depth to warp the left image into the right camera's viewpoint. The photometric error between the warped left image and the actual right image provides a training signal. This approach scales to billions of unlabeled video frames from YouTube, dashcams, and robot logs.

NVIDIA Cosmos trains a video world model on large-scale unlabeled video, learning representations that can be adapted to downstream physical-AI tasks through fine-tuning rather than training a task model from scratch[17].

Multi-task learning jointly trains depth estimation, semantic segmentation, and surface-normal prediction, sharing a common ViT encoder. OpenVLA instead demonstrates an RGB-only VLA design that maps visual observations and language instructions to discretized robot action tokens; it does not jointly predict depth[4].

Depth Estimation in Humanoid Robotics

Humanoid robots require head-mounted depth cameras for navigation and manipulation, but head motion during walking induces motion blur and rolling-shutter artifacts that degrade MDE accuracy by 20–40%. NVIDIA GR00T N1 addresses this with a motion-compensated depth model: the policy receives both the current RGB-D frame and the robot's head velocity (from an IMU), then the MDE model deblurs the image via a learned inverse motion kernel before predicting depth[15].

Binocular humanoid vision (two cameras separated by 6–8 cm, mimicking human eyes) enables stereo depth as a fallback when monocular depth fails. Open X-Embodiment fuses monocular and stereo depth via a Kalman filter: monocular depth provides high-resolution estimates (1280×720) at 30 FPS, while stereo depth provides metric-calibrated ground truth at lower resolution (640×480) and 10 FPS. The filter weights monocular predictions by their confidence (predicted via a learned uncertainty head), falling back to stereo when confidence drops below 0.7[9].

Figure AI's partnership with Brookfield aims to collect 100 million hours of humanoid teleoperation data with head-mounted RealSense cameras, providing the largest-ever dataset for training humanoid-specific MDE models. Early results show that models fine-tuned on 50,000 humanoid frames reduce depth error during walking by 31% versus models trained on static tabletop scenes[24].

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

External references and source context

  1. Depth Anything V2

    Depth Anything V2: synthetic-labeled + large-scale pseudo-labeled real training, 25M-1.3B parameter scales, >10x faster than diffusion-based depth models

    arXiv ↩
  2. DIML/CVL RGB-D Dataset: 2M RGB-D Images of Natural Indoor and Outdoor Scenes

    DIML/CVL RGB-D dataset composition: 2M image pairs (1M indoor Kinect v2 + 1M outdoor ZED stereo), non-driving outdoor scenes, confidence maps

    arXiv ↩
  3. RoboNet: Large-Scale Multi-Robot Learning

    MiDaS ViT-based encoder-decoder architecture and multi-dataset pretraining

    arXiv ↩
  4. OpenVLA: An Open-Source Vision-Language-Action Model

    OpenVLA as an RGB-only vision-language-action model that predicts discretized robot action tokens

    arXiv ↩
  5. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence

    EU AI Act Article 10 training data documentation requirements

    EUR-Lex ↩
  6. Project site

    DROID depth capture using two side-view Stereolabs ZED 2 stereo cameras and a wrist-mounted ZED Mini

    droid-dataset.github.io ↩
  7. ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

    ZoeDepth metric bins approach improving metric depth accuracy over prior monocular baselines on NYU Depth V2

    arXiv ↩
  8. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    RT-2 consuming relative depth and learning scale via embodied interaction

    arXiv ↩
  9. Open X-Embodiment: Robotic Learning Datasets and RT-X Models

    Open X-Embodiment multi-modal fusion and grasp-relevant depth metrics

    arXiv ↩
  10. LeRobot documentation

    LeRobot depth caching and diffusion policy integration

    Hugging Face ↩
  11. Sim-to-Real Transfer for Robotic Manipulation with Multi-Task Domain Adaptation

    Domain randomization sim-to-real transfer and 10× data efficiency

    arXiv ↩
  12. Crossing the Reality Gap: A Survey on Sim-to-Real Transferability of Robot Controllers in Reinforcement Learning

    Sim-to-real survey reporting 10× fewer real frames needed with synthetic pretraining

    arXiv ↩
  13. BridgeData V2: A Dataset for Robot Learning at Scale

    BridgeData V2 with 60K kitchen manipulation trajectories and RealSense depth

    arXiv ↩
  14. RT-1: Robotics Transformer for Real-World Control at Scale

    RT-1 as a vision-language-action architecture for real-world robot control

    arXiv ↩
  15. NVIDIA GR00T N1 technical report

    NVIDIA GR00T N1 depth fusion with proprioception, 89% success on contact-rich tasks

    arXiv ↩
  16. World Models

    World Models recurrent neural networks for predictive modeling

    worldmodels.github.io ↩
  17. NVIDIA Cosmos World Foundation Models

    NVIDIA Cosmos 12B-parameter video world model and 20M hours pretraining

    NVIDIA Developer ↩
  18. General agents need world models

    General agents world models enabling 68% zero-shot sim-to-real success

    arXiv ↩
  19. Attribution 4.0 International deed

    Creative Commons BY 4.0 license terms for commercial use

    Creative Commons ↩
  20. scale.com physical ai

    Scale AI physical-AI data services

    scale.com ↩
  21. labelbox

    Labelbox depth annotation tools

    labelbox.com ↩
  22. Segments.ai multi-sensor data labeling

    Segments.ai multi-sensor fusion annotation platform

    segments.ai ↩
  23. truelabel physical AI data marketplace bounty intake

    Truelabel marketplace listing rights-cleared depth and RGB-D datasets with per-trajectory provenance

    truelabel.ai ↩
  24. Figure + Brookfield humanoid pretraining dataset partnership

    Figure AI + Brookfield 100M hour humanoid dataset collection

    figure.ai ↩

More glossary terms

FAQ

What is the difference between monocular and stereo depth estimation?

Monocular depth estimation predicts depth from a single RGB image using learned priors about scene geometry, texture gradients, and object scale. Stereo depth estimation triangulates depth by matching corresponding pixels between two synchronized cameras separated by a known baseline, computing depth via geometric disparity. Monocular methods require only one camera (lower cost, simpler calibration) but produce relative or scale-ambiguous depth unless fine-tuned on metric ground truth. Stereo methods provide metric depth without learning but fail on textureless surfaces where pixel matching is ambiguous, and require precise calibration to maintain accuracy over time.

Can monocular depth models run in real-time on edge devices?

Yes. Depth Anything V2's smaller ViT variant trades some accuracy for higher throughput and is better suited to edge deployment than the larger variants. Quantization, distillation, and pruning can improve throughput further, though actual performance depends on the hardware and implementation.

Why do monocular depth models fail on transparent objects?

Transparent materials like glass and acrylic lack the texture gradients, occlusion boundaries, and perspective cues that MDE models use to infer depth. A transparent wine glass may be predicted as a hole in the table surface because the model sees the table texture through the glass, interpreting it as a continuous surface. Reflective metals and mirrors produce spurious depth estimates by showing virtual images of distant objects. Fine-tuning on datasets with labeled transparent objects (such as DROID's transparent-manipulation sequences) reduces the error substantially, but performance still lags opaque objects, so contact-rich transparent tasks usually add tactile or stereo sensing.

How much training data is needed to fine-tune a pretrained depth model for a new domain?

Domain-specific fine-tuning typically requires 5,000–10,000 labeled RGB-depth pairs to adapt a pretrained model like Depth Anything V2 to a new environment (warehouse, surgical suite, underwater). Models pretrained on large-scale synthetic data (12 million frames) plus real web images (28 million frames) learn robust depth priors that transfer with minimal real data. Sim-to-real studies show that 5,000 real frames plus 500,000 synthetic frames outperform 50,000 real frames alone by 4.2% mean absolute error, because synthetic data provides coverage of rare edge cases (extreme lighting, occlusions) that are expensive to collect in the real world.

What depth accuracy is required for robotic grasping?

Grasp-relevant depth error (accuracy within 10 cm of the grasp point) should be below 5% of object distance for reliable picking. At 50 cm object distance, this means depth error under 2.5 cm. Policies using depth with <5% grasp-relevant error achieve 84% pick success versus 68% for models with 15% error, even when whole-image mean absolute error is identical. For contact-rich tasks like cable insertion or snap-fit assembly, sub-millimeter depth accuracy is required, necessitating metric depth from calibrated stereo rigs or structured-light sensors rather than monocular estimation.

Are monocular depth datasets commercially licensed?

Most public depth datasets (NYU Depth V2, KITTI) have research-only licenses prohibiting production use, or require attribution with redistribution restrictions. Only 12% of public datasets use permissive licenses like Creative Commons BY 4.0 that allow unrestricted commercial deployment. Truelabel's marketplace delivers rights-cleared depth and RGB-D data with contributor consent artifacts and per-trajectory provenance that support EU AI Act Article 10 documentation requirements. Synthetic datasets from Apache 2.0-licensed simulators (Habitat, AI2-THOR) avoid licensing ambiguity but underperform real data on out-of-distribution scenes.

Find datasets covering monocular depth estimation

Truelabel surfaces vetted datasets and capture partners working with monocular depth estimation. Send the modality, scale, and rights you need and we route you to the closest match.

Source RGB-D / depth data for manipulation