Monocular Depth Estimation on DIODE
3.14AbsRelMoGe-2
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MoGe-2Method Type=Discriminative depth estimation2026.06 | 3.14 | 97.4 | — | |
| MoGeFine-tuning=false, Evaluation Protocol=Scale-invariant relative depth2025.11 | 3.23 | 97.4 | — | |
| Modality ForcingMethod Type=I2D from Joint Model2026.06 | 3.35 | 97.7 | — | |
| VGGT+OursFine-tuning=true, Evaluation Protocol=Scale-invariant relative depth2025.11 | 3.59 | 96.7 | — | |
| UniDepth V2Method Type=Discriminative depth estimation2026.06 | 4.05 | 96.5 | — | |
| MoGeAlignment=Optimized scale and shift factor2026.03 | 4.37 | 96.4 | — | |
| Depth ProMethod Type=Discriminative depth estimation2026.06 | 4.66 | 95.6 | — | |
| DUSt3R+OursFine-tuning=true, Evaluation Protocol=Scale-invariant relative depth2025.11 | 4.7 | 95.3 | — | |
| MoGe 2Type=Discriminative, Evaluation Protocol=Zero-shot2026.01 | 4.8 | 97.1 | — | |
| DAGEAlignment=Optimized scale and shift factor2026.03 | 4.97 | 94.6 | — | |
| PPDMethod Type=Generative depth estimation2026.06 | 4.97 | 95.6 | — | |
| CUT3R+OursFine-tuning=true, Evaluation Protocol=Scale-invariant relative depth2025.11 | 5.08 | 94.8 | — | |
| MoGe2Alignment=Optimized scale and shift factor2026.03 | 5.13 | 94.9 | — | |
| PPDType=Generative, Evaluation Protocol=Zero-shot2026.01 | 5.2 | 97 | — | |
| VGGTFine-tuning=false, Evaluation Protocol=Scale-invariant relative depth2025.11 | 5.24 | 94.5 | — | |
| DA-v2Method Type=Discriminative depth estimation2026.06 | 5.41 | 94.6 | — | |
| MASt3RMethod Type=Discriminative depth estimation2026.06 | 5.79 | 94.1 | — | |
| CUT3RFine-tuning=false, Evaluation Protocol=Scale-invariant relative depth2025.11 | 5.93 | 93.2 | — | |
| Depth ProType=Discriminative, Evaluation Protocol=Zero-shot2026.01 | 6.1 | 95.9 | — | |
| MarigoldMethod Type=Generative depth estimation2026.06 | 6.13 | 94.5 | — | |
| DepthProAlignment=Optimized scale and shift factor2026.03 | 6.28 | 93.7 | — | |
| GeoWizardMethod Type=Generative depth estimation2026.06 | 6.37 | 94 | — | |
| DepthAnything v2Prior=DINOv2 - 142M, training samples=63M2025.12 | 6.5 | 95.4 | — | |
| LotusMethod Type=Generative depth estimation2026.06 | 6.7 | 93.8 | — | |
| DUSt3RFine-tuning=false, Evaluation Protocol=Scale-invariant relative depth2025.11 | 6.85 | 92.4 | — | |
| ZoeDepthMethod Type=Discriminative depth estimation2026.06 | 7.8 | 90.9 | — | |
| Depth Anything v2Type=Discriminative, Evaluation Protocol=Zero-shot2026.01 | 8 | 95.2 | — | |
| Depth Prowith Ours=true2026.05 | 9.3 | 88.9 | — | |
| DepthAnythingV2with Ours=true2026.05 | 9.3 | 91.7 | — | |
| Marigoldwith Ours=true2026.05 | 9.4 | 89.6 | — | |
| DepthMasterwith Ours=true2026.05 | 9.5 | 88.8 | — | |
| DPTwith Ours=true2026.05 | 9.6 | 88.7 | — | |
| UniDepthV2with Ours=true2026.05 | 9.6 | 88.7 | — | |
| LeReSwith Ours=true2026.05 | 9.7 | 88.5 | — | |
| GeoWizardwith Ours=true2026.05 | 9.7 | 88.5 | — | |
| LotusType=Generative, Evaluation Protocol=Zero-shot2026.01 | 9.8 | 92.4 | — | |
| Metric3Dv2with Ours=true2026.05 | 9.9 | 88.1 | — | |
| MarigoldType=Generative, Evaluation Protocol=Zero-shot2026.01 | 10 | 90.7 | — | |
| JointDiTMethod Type=I2D from Joint Model2026.06 | 10.22 | 93.9 | — | |
| Lotuswith Ours=true2026.05 | 10.8 | 86.8 | — | |
| MiDaSwith Ours=true2026.05 | 11.3 | 86.4 | — | |
| GeoWizardType=Generative, Evaluation Protocol=Zero-shot2026.01 | 12 | 89.8 | — | |
| UniConMethod Type=I2D from Joint Model2026.06 | 13.57 | 90.8 | — | |
| Student-PointmapCamera Intrinsics=Unknown2026.01 | 13.9 | 79.8 | — | |
| Student-PointMapCamera Intrinsics=Provided2026.01 | 14.1 | 74.3 | — | |
| Cross-Context DistillationZero-shot=true, Teacher Model=MiDaS v3.12025.02 | 14.2 | 78.8 | — | |
| MoGe-2Fine-tuning=false, Evaluation Protocol=Metric depth2025.11 | 15.97 | 71.3 | — | |
| Metric3Dv2GT aligned=false2024.11 | 16 | 88 | — | |
| MoGe-2Camera Intrinsics=Provided2026.01 | 16.2 | 77.1 | — | |
| UniDepth V1Camera Intrinsics=Unknown2026.01 | 17.1 | 71.9 | — | |
| MoGe-2Camera Intrinsics=Unknown2026.01 | 17.5 | 66.4 | — | |
| DepthLabSynthetic Training Samples=74k2024.12 | 17.6 | 85.6 | — | |
| DPTReal Training Samples=1.2M, Synthetic Training Samples=188K2024.12 | 18.1 | 75.8 | — | |
| DPTZero-shot=true2025.02 | 18.2 | 75.8 | — | |
| DPTPrior=ImageNet - 14M, training samples=1.4M2025.12 | 18.2 | 75.8 | — | |
| JointNetMethod Type=I2D from Joint Model2026.06 | 20.02 | 82.7 | — | |
| Omnizero-shot=true2026.04 | 20.34 | 83.83 | — | |
| DA3 giantzero-shot=true2026.04 | 20.5 | 82.69 | — | |
| UniDepth V1Camera Intrinsics=Provided2026.01 | 21 | 63.5 | — | |
| VGGTzero-shot=true2026.04 | 21.15 | 82.15 | — | |
| DepthMasterTraining Data=74K, zero-shot=true, affine-invariant=true2025.01 | 21.5 | 77.6 | — | |
| DepthMasterwith Ours=false2026.05 | 21.5 | 77.6 | — | |
| BetterDepth-400Training samples=400, Feed-forward=true, Diffusion=true, Zero-shot=true2024.07 | 21.9 | 75.3 | — | |
| BetterDepth-2KTraining samples=2K, Feed-forward=true, Diffusion=true, Zero-shot=true2024.07 | 22 | 75.5 | — | |
| VAR-DepthPrior=Switti - 100M, training samples=74K, config=optimized wk2025.12 | 22.3 | 75.4 | — | |
| DepthFMTraining Data=74K, zero-shot=true, affine-invariant=true2025.01 | 22.4 | 78.5 | — | |
| DepthFMSynthetic Training Samples=63K2024.12 | 22.4 | 79.8 | — | |
| DepthFMTraining samples=63K, Feed-forward=false, Diffusion=true, Zero-shot=true2024.07 | 22.5 | 80 | — | |
| DepthFMPrior=SD - 2.3B, training samples=63K2025.12 | 22.5 | 80 | — | |
| BetterDepthTraining samples=74K, Feed-forward=true, Diffusion=true, Zero-shot=true2024.07 | 22.6 | 75.5 | — | |
| GenPerceptZero-shot=true2025.02 | 22.6 | 74.1 | — | |
| Marigoldzero-shot=true2026.04 | 22.66 | 81.64 | — | |
| LotusTraining Data=59K, zero-shot=true, affine-invariant=true2025.01 | 22.8 | 73.8 | — | |
| Lotuswith Ours=false2026.05 | 22.8 | 73.8 | — | |
| JetViT-DepthAnythingSize(Spec)=Large(0 FA), Latency (ms)=29.70, Throughput (samples/s)=34.97, Zero-shot evaluation protocol=true2026.05 | 22.8 | 74.3 | — | |
| JetViT-DepthAnythingSize(Spec)=Large(2 FA), Latency (ms)=32.63, Throughput (samples/s)=32.13, Zero-shot evaluation protocol=true2026.05 | 22.8 | 74.9 | — | |
| MASt3R+OursFine-tuning=true, Evaluation Protocol=Metric depth2025.11 | 22.84 | 50.1 | — | |
| JetViT-DepthAnythingSize(Spec)=Large(1 FA), Latency (ms)=31.04, Throughput (samples/s)=33.72, Zero-shot evaluation protocol=true2026.05 | 22.9 | 74.4 | — | |
| DepthAnythingV2Size(Spec)=Large, Latency (ms)=60.76, Throughput (samples/s)=14.20, Zero-shot evaluation protocol=true2026.05 | 23.1 | 74.9 | — | |
| JetViT-DepthAnythingSize(Spec)=Giant(2 FA), Latency (ms)=90.65, Throughput (samples/s)=11.39, Zero-shot evaluation protocol=true2026.05 | 23.1 | 75 | — | |
| Cross-Context DistillationZero-shot=true, Teacher Model=DepthAnythingv2-Large2025.02 | 23.3 | 75.3 | — | |
| DepthAnythingV2Size(Spec)=Giant, Latency (ms)=164.25, Throughput (samples/s)=6.35, Zero-shot evaluation protocol=true2026.05 | 23.4 | 75.1 | — | |
| DAV3-Metric-LargeCamera Intrinsics=Unknown2026.01 | 24.2 | 53.7 | — | |
| HDNReal Training Samples=300K2024.12 | 24.2 | 78.3 | — | |
| HDNTraining samples=300K, Feed-forward=true, Diffusion=false, Zero-shot=true2024.07 | 24.6 | 78 | — | |
| HDNTraining Data=300K, zero-shot=true, affine-invariant=true2025.01 | 24.6 | 78 | — | |
| HDNZero-shot=true2025.02 | 24.6 | 78 | — | |
| HDNPrior=ImageNet - 14M, training samples=300K2025.12 | 24.6 | 78 | — | |
| Metric3Dv2with Ours=false2026.05 | 24.6 | 82.3 | — | |
| UniDepthV2with Ours=false2026.05 | 25.1 | 79.5 | — | |
| Depth AnythingTraining samples=63.5M, Feed-forward=true, Diffusion=false, Zero-shot=true2024.07 | 26 | 75.9 | — | |
| UniDepthGT aligned=false2024.11 | 26 | 66 | — | |
| Depth Anything V2Training Data=63.5M, zero-shot=true, affine-invariant=true2025.01 | 26 | 75.9 | — | |
| DepthAnythingV2with Ours=false2026.05 | 26 | 75.9 | — | |
| DepthAnything v2Zero-shot=true2025.02 | 26.2 | 75.4 | — | |
| MiDaS V3 DPTSize(Spec)=Large(Swin), Latency (ms)=39.28, Throughput (samples/s)=27.14, Zero-shot evaluation protocol=true2026.05 | 26.2 | 69.6 | — | |
| MiDaSTraining samples=2M, Feed-forward=true, Diffusion=false, Zero-shot=true2024.07 | 26.6 | 71.3 | — | |
| MiDaSTraining Data=2M, zero-shot=true, affine-invariant=true2025.01 | 26.6 | 71.3 | — | |
| MiDaSwith Ours=false2026.05 | 26.6 | 71.3 | — | |
| DPTTraining samples=1.4M, Feed-forward=true, Diffusion=false, Zero-shot=true2024.07 | 26.9 | 73 | — |