Monocular Depth Estimation on DDAD (test)
5.307RMSEGRIN
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GRINzero_shot=true, require_camera_intrinsics=true2024.09 | 5.307 | 0.093 | — | — | 92.2 | — | — | — | — | — | — | — | — | — | |
| DMDzero_shot=true, require_camera_intrinsics=true2024.09 | 5.365 | 0.108 | — | — | 90.7 | — | — | — | — | — | — | — | — | — | |
| UniDepthzero_shot=true, model_variant=UniDepth-C, require_camera_intrinsics=true2024.09 | 5.399 | 0.097 | — | — | 91.9 | — | — | — | — | — | — | — | — | — | |
| ScaleDepth-NKtrained on extra depth datasets=false, zero-shot generalization=true, max evaluation depth=80m2024.07 | 6.097 | 0.121 | — | — | — | — | — | — | — | — | 87.1 | — | — | — | |
| NeWCRFsZero-shot=true2023.02 | 6.183 | 0.119 | — | — | 87.4 | 0 | — | — | — | — | — | — | — | — | |
| NeWCRFstrained on extra depth datasets=false, zero-shot generalization=true, max evaluation depth=80m2024.07 | 6.183 | 0.119 | — | — | — | — | — | — | — | — | 87.4 | — | — | — | |
| NeWCRFszero_shot=true, require_camera_intrinsics=true2024.09 | 6.183 | 0.119 | — | — | 87.4 | — | — | — | — | — | — | — | — | — | |
| NeWCRFSScaling=-, Training Dataset=KITTI2024.10 | 6.183 | 0.119 | — | — | 87.4 | — | — | — | — | — | — | — | — | — | |
| ZeroDepthzero_shot=true, require_camera_intrinsics=true2024.09 | 6.318 | 0.1 | — | — | 88.9 | — | — | — | — | — | — | — | — | — | |
| ScaleDepth-Ktrained on extra depth datasets=false, zero-shot generalization=true, max evaluation depth=80m2024.07 | 6.378 | 0.12 | — | — | — | — | — | — | — | — | 86.3 | — | — | — | |
| ZoeD-M12-KZero-shot=true, Fine-tuning dataset=KITTI, Backbone=MiDaS v3.1 (M12)2023.02 | 7.108 | 0.129 | — | — | 83.5 | -9.3 | — | — | — | — | — | — | — | — | |
| ZoeDepth-M12Scaling=Image, Training Dataset=KITTI2024.10 | 7.108 | 0.129 | — | — | 83.5 | — | — | — | — | — | — | — | — | — | |
| ZoeD-M12-NKZero-shot=true, Fine-tuning dataset=NYU Depth v2 + KITTI, Backbone=MiDaS v3.1 (M12)2023.02 | 7.225 | 0.138 | — | — | 82.4 | -12.8 | — | — | — | — | — | — | — | — | |
| ZoeD-M12-NKtrained on extra depth datasets=true, zero-shot generalization=true, max evaluation depth=80m2024.07 | 7.225 | 0.138 | — | — | — | — | — | — | — | — | 82.4 | — | — | — | |
| ZoeDepthzero_shot=true, require_camera_intrinsics=false2024.09 | 7.225 | 0.138 | — | — | 82.4 | — | — | — | — | — | — | — | — | — | |
| ZoeDepth-M12Scaling=Image, Training Dataset=NYUv2, KITTI2024.10 | 7.225 | 0.138 | — | — | 82.4 | — | — | — | — | — | — | — | — | — | |
| AFNetType=Multi View2024.03 | 7.23 | 0.088 | 1.41 | — | — | — | — | — | — | — | — | — | — | — | |
| BTSZero-shot=true2023.02 | 7.55 | 0.147 | — | — | 80.5 | -17.8 | — | — | — | — | — | — | — | — | |
| BTStrained on extra depth datasets=false, zero-shot generalization=true, max evaluation depth=80m2024.07 | 7.55 | 0.147 | — | — | — | — | — | — | — | — | 80.5 | — | — | — | |
| AdaBinszero_shot=true, require_camera_intrinsics=true2024.09 | 7.55 | 0.147 | — | — | 76.6 | — | — | — | — | — | — | — | — | — | |
| ZoeD-X-KZero-shot=true, Fine-tuning dataset=KITTI2023.02 | 7.734 | 0.137 | — | — | 79 | -16.6 | — | — | — | — | — | — | — | — | |
| ZoeD-X-Ktrained on extra depth datasets=false, zero-shot generalization=true, max evaluation depth=80m2024.07 | 7.734 | 0.137 | — | — | — | — | — | — | — | — | 79 | — | — | — | |
| ZoeDepth-XScaling=Image, Training Dataset=KITTI2024.10 | 7.734 | 0.137 | — | — | 89 | — | — | — | — | — | — | — | — | — | |
| IterMVSType=Multi View2024.03 | 7.95 | 0.104 | 1.59 | — | — | — | — | — | — | — | — | — | — | — | |
| LocalBinsZero-shot=true2023.02 | 8.139 | 0.151 | — | — | 77.7 | -23.2 | — | — | — | — | — | — | — | — | |
| MVSNetType=Multi View2024.03 | 8.21 | 0.109 | 1.62 | — | — | — | — | — | — | — | — | — | — | — | |
| AdaBinsZero-shot=true2023.02 | 8.56 | 0.154 | — | — | 76.6 | -26.7 | — | — | — | — | — | — | — | — | |
| AdaBinstrained on extra depth datasets=false, zero-shot generalization=true, max evaluation depth=80m2024.07 | 8.56 | 0.154 | — | — | — | — | — | — | — | — | 76.6 | — | — | — | |
| AdabinsScaling=-, Training Dataset=KITTI2024.10 | 8.56 | 0.154 | — | — | 79 | — | — | — | — | — | — | — | — | — | |
| Embodied Depth Estimation2025.03 | 8.673 | 0.145 | — | — | 82.3 | — | — | — | — | — | — | 94.7 | 98 | — | |
| iDisc2025.03 | 8.989 | 0.163 | — | — | 80.9 | — | — | — | — | — | — | 93.4 | 97.1 | — | |
| MaGNetType=Multi View2024.03 | 9.23 | 0.112 | 1.74 | — | — | — | — | — | — | — | — | — | — | — | |
| DepthAnything (finetuned)Supervision=Supervised, Evaluation Protocol=finetuned, Params.=335.79 M, Latency=142.26 ms2025.12 | 9.475 | 0.14 | 1.866 | 0.228 | 83.1 | — | — | — | — | — | — | 93.5 | 96.9 | — | |
| MVS2DType=Multi View2024.03 | 9.82 | 0.132 | 2.05 | — | — | — | — | — | — | — | — | — | — | — | |
| CasMVSType=Multi View2024.03 | 9.87 | 0.129 | 2.01 | — | — | — | — | — | — | — | — | — | — | — | |
| FutureDepthEncoder=Swin-L, Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 10.016 | — | 2.96 | — | 83.3 | — | — | — | — | — | — | — | — | — | |
| Adabins2025.03 | 10.24 | 0.201 | — | — | 74.8 | — | — | — | — | — | — | 91.2 | 96.2 | — | |
| DPT (Linear Fit)Backbone=DPT, Scaling=Linear Fit, Training Dataset=DDAD2024.10 | 10.342 | 0.163 | — | 0.254 | 80.2 | — | — | — | — | — | — | 0.954 | 0.99 | — | |
| BinsFormerGround embedding module=GE-Adaptive2023.09 | 10.459 | 0.145 | 2.101 | 0.235 | — | — | 22.06 | 65.9 | 50.4 | — | — | — | — | — | |
| NeWCRFs + MAMOEncoder=Swin-large, Max distance=200 meters, Input frame resolution=1216 x 1936, Training dataset=KITTI2023.07 | 10.462 | — | 2.99 | — | 86.7 | — | — | — | — | — | — | — | — | — | |
| MonoViT + MB + OursSupervision=Supervised, Median Scaling=false, Params.=28.31 M, Latency=27.71 ms2025.12 | 10.483 | 0.166 | 2.26 | 0.256 | 77.3 | — | — | — | — | — | — | 91.6 | 96.1 | 7.54 | |
| PMNetType=Multi View2024.03 | 10.56 | 0.141 | 2.23 | — | — | — | — | — | — | — | — | — | — | — | |
| BinsFormerGround embedding module=GE-Vanilla2023.09 | 10.561 | 0.146 | 2.109 | 0.235 | — | — | 22.252 | 65.8 | 50.2 | — | — | — | — | — | |
| DepthFormerGround embedding module=GE-Adaptive2023.09 | 10.596 | 0.145 | 2.119 | 0.237 | — | — | 22.19 | 65.6 | 50 | — | — | — | — | — | |
| DepthFormerGround embedding module=GE-Vanilla2023.09 | 10.79 | 0.149 | 2.121 | 0.24 | — | — | 22.437 | 65 | 49.3 | — | — | — | — | — | |
| PixelFormerGround embedding module=GE-Adaptive2023.09 | 10.803 | 0.145 | 2.122 | 0.241 | — | — | 22.268 | 66.1 | 50.5 | — | — | — | — | — | |
| PixelFormerGround embedding module=GE-Vanilla2023.09 | 10.848 | 0.148 | 2.123 | 0.241 | — | — | 22.272 | 66 | 50.4 | — | — | — | — | — | |
| BinsFormer2023.09 | 10.866 | 0.149 | 2.142 | 0.244 | — | — | 22.513 | 65.3 | 49.6 | — | — | — | — | — | |
| BinsFormer2025.03 | 10.866 | 0.149 | — | — | 73.2 | — | — | — | — | — | — | 89.4 | 95.6 | — | |
| PixelFormer2023.09 | 10.92 | 0.151 | 2.14 | 0.242 | — | — | 22.311 | 65.9 | 50.2 | — | — | — | — | — | |
| PixelFormer2025.03 | 10.92 | 0.151 | — | — | 69.1 | — | — | — | — | — | — | 85.6 | 94.9 | — | |
| NeWCRF2025.03 | 10.98 | 0.219 | — | — | 70.2 | — | — | — | — | — | — | 88.1 | 95.1 | — | |
| DepthFormer2023.09 | 11.051 | 0.152 | 2.23 | 0.246 | — | — | 22.629 | 64.9 | 49.3 | — | — | — | — | — | |
| DepthFormer2025.03 | 11.051 | 0.152 | — | — | 68.9 | — | — | — | — | — | — | 84.5 | 94.7 | — | |
| AdaBinsType=Single View2024.03 | 11.08 | 0.164 | 2.66 | — | — | — | — | — | — | — | — | — | — | — | |
| PixelFormer + MAMoEncoder=Swin-large, Max distance=200 meters, Input frame resolution=1216 x 1936, Training dataset=KITTI2023.07 | 11.094 | — | 3.349 | — | 87 | — | — | — | — | — | — | — | — | — | |
| MAMOEncoder=Swin-L, Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 11.094 | — | 3.349 | — | 87 | — | — | — | — | — | — | — | — | — | |
| BTSGround embedding module=GE-Adaptive2023.09 | 11.186 | 0.156 | 2.36 | 0.253 | — | — | 23.761 | 64.2 | 48.5 | — | — | — | — | — | |
| BTSGround embedding module=GE-Vanilla2023.09 | 11.219 | 0.158 | 2.377 | 0.253 | — | — | 23.863 | 63.6 | 47.9 | — | — | — | — | — | |
| BTS2023.09 | 11.466 | 0.162 | 2.492 | 0.259 | — | — | 24.314 | 63.6 | 47.8 | — | — | — | — | — | |
| BTS2025.03 | 11.466 | 0.162 | — | — | 75.7 | — | — | — | — | — | — | 91.3 | 96.2 | — | |
| HRDepth + MB + OursSupervision=Supervised, Median Scaling=false, Params.=14.8 M, Latency=15.18 ms2025.12 | 11.578 | 0.183 | 2.684 | 0.283 | 73.6 | — | — | — | — | — | — | 89.6 | 95 | 5.47 | |
| SwinV2-MIMEncoder=Swin-large, Max distance=200 meters, Input frame resolution=1216 x 1936, Training dataset=KITTI2023.07 | 11.641 | — | 3.505 | — | 85.3 | — | — | — | — | — | — | — | — | — | |
| MonoViT + MBSupervision=Supervised, Median Scaling=false, Params.=28.31 M, Latency=27.71 ms2025.12 | 11.757 | 0.178 | 2.585 | 0.286 | 72.8 | — | — | — | — | — | — | 89.3 | 95 | 0.34 | |
| MonoViTSupervision=Supervised, Median Scaling=false, Params.=27.87 M, Latency=18.83 ms2025.12 | 11.777 | 0.179 | 2.602 | 0.287 | 72.5 | — | — | — | — | — | — | 89.2 | 94.9 | 0 | |
| BTSType=Single View2024.03 | 11.85 | 0.169 | 2.81 | — | — | — | — | — | — | — | — | — | — | — | |
| DepthAnything (zero-shot)Supervision=Supervised, Evaluation Protocol=zero-shot, Params.=335.79 M, Latency=142.26 ms2025.12 | 11.866 | 0.27 | 3.291 | 0.621 | 60.1 | — | — | — | — | — | — | 82.6 | 90.4 | — | |
| EGA-Depth-MRSupervision=Self-supervised, Median Scaling=true, Params.=not public, Latency=not public2025.12 | 11.922 | 0.191 | 3.126 | 0.29 | 74.7 | — | — | — | — | — | — | 90.1 | 95 | — | |
| PackNet-SAN2023.09 | 11.936 | 0.187 | 2.776 | 0.276 | — | — | — | — | — | — | — | — | — | — | |
| PackNet-SAN2025.03 | 11.936 | 0.187 | — | — | 68.4 | — | — | — | — | — | — | 82.1 | 93.2 | — | |
| NeWCRFsEncoder=Swin-large, Max distance=200 meters, Input frame resolution=1216 x 1936, Training dataset=KITTI2023.07 | 11.956 | — | 4.041 | — | 81.6 | — | — | — | — | — | — | — | — | — | |
| NeWCRFsEncoder=Swin-L, Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 11.956 | — | 4.041 | — | 81.6 | — | — | — | — | — | — | — | — | — | |
| EGA-Depth-LRSupervision=Self-supervised, Median Scaling=true, Params.=not public, Latency=not public2025.12 | 12.117 | 0.195 | 3.211 | 0.297 | 74.3 | — | — | — | — | — | — | 89.6 | 94.7 | — | |
| Monodepth2 + MB + OursSupervision=Supervised, Median Scaling=false, Params.=15.03 M, Latency=12.77 ms2025.12 | 12.134 | 0.191 | 2.865 | 0.3 | 71 | — | — | — | — | — | — | 88.1 | 94.3 | 4.64 | |
| Metric3DType=Single View2024.03 | 12.15 | 0.183 | 2.92 | — | — | — | — | — | — | — | — | — | — | — | |
| SurroundDepthSupervision=Self-supervised, Median Scaling=true, Params.=59.56 M, Latency=11.79 ms2025.12 | 12.27 | 0.2 | 3.392 | 0.301 | 74 | — | — | — | — | — | — | 89.4 | 94.7 | — | |
| HRDepth + MBSupervision=Supervised, Median Scaling=false, Params.=14.8 M, Latency=15.18 ms2025.12 | 12.399 | 0.194 | 2.961 | 0.303 | 70 | — | — | — | — | — | — | 87.7 | 94.2 | 0.28 | |
| HRDepthSupervision=Supervised, Median Scaling=false, Params.=14.1 M, Latency=6.19 ms2025.12 | 12.428 | 0.196 | 2.955 | 0.304 | 69.9 | — | — | — | — | — | — | 87.5 | 94 | 0 | |
| RSABackbone=DPT, Scaling=RSA (Ours), Training Dataset=NYUv2, KITTI, VOID2024.10 | 12.437 | 0.165 | — | 0.276 | 76.8 | — | — | — | — | — | — | 0.942 | 0.983 | — | |
| FeatDepthType=Single View2024.03 | 12.45 | 0.189 | 3.21 | — | — | — | — | — | — | — | — | — | — | — | |
| PixelFormerEncoder=Swin-large, Max distance=200 meters, Input frame resolution=1216 x 1936, Training dataset=KITTI2023.07 | 12.467 | — | 4.474 | — | 80.2 | — | — | — | — | — | — | — | — | — | |
| PixelFormerEncoder=Swin-L, Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 12.467 | — | 4.474 | — | 80.2 | — | — | — | — | — | — | — | — | — | |
| Monodepth2 + MBSupervision=Supervised, Median Scaling=false, Params.=15.03 M, Latency=12.77 ms2025.12 | 12.727 | 0.195 | 3.016 | 0.322 | 68.6 | — | — | — | — | — | — | 86.5 | 93.2 | 1.08 | |
| BaselineEncoder=Swin-L, Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 12.841 | — | 4.506 | — | 80.4 | — | — | — | — | — | — | — | — | — | |
| Monodepth2Supervision=Supervised, Median Scaling=false, Params.=14.33 M, Latency=3.75 ms2025.12 | 12.849 | 0.2 | 3.087 | 0.323 | 67.9 | — | — | — | — | — | — | 86.1 | 93.2 | 0 | |
| Monodepth2Supervision=Self-supervised, Median Scaling=true, Params.=14.33 M, Latency=3.75 ms2025.12 | 12.962 | 0.217 | 3.641 | 0.323 | 69.9 | — | — | — | — | — | — | 87.7 | 93.9 | — | |
| PackNet-SfMSupervision=Self-supervised, Median Scaling=true, Params.=128.29 M, Latency=63.08 ms2025.12 | 13.253 | 0.234 | 3.802 | 0.331 | 67.2 | — | — | — | — | — | — | 86 | 93.1 | — | |
| Monodepth2Type=Single View2024.03 | 13.32 | 0.191 | 3.52 | — | — | — | — | — | — | — | — | — | — | — | |
| PackNet-SfMBackbone=PackNet, ImageNet Pre-training=false, Resolution=640 x 384, Distance Range=0-200m2019.05 | 13.452 | 0.162 | 3.917 | 0.269 | 82.3 | — | — | — | — | — | — | — | — | — | |
| RSABackbone=DPT, Scaling=RSA (Ours), Training Dataset=NYUv2, KITTI2024.10 | 13.539 | 0.171 | — | 0.284 | 77.7 | — | — | — | — | — | — | 0.938 | 0.981 | — | |
| ManyDepth-FSEncoder=Swin-large, Max distance=200 meters, Input frame resolution=1216 x 1936, Training dataset=KITTI2023.07 | 13.899 | — | 4.211 | — | 78.4 | — | — | — | — | — | — | — | — | — | |
| ManyDepth-FSEncoder=Swin-L, Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 13.899 | — | 4.211 | — | 78.4 | — | — | — | — | — | — | — | — | — | |
| iDiscTrain set=KITTI Eigen-split, Evaluation mode=Zero-shot, Fine-tuning=None2023.04 | 14.26 | 0.367 | — | — | 35 | — | 29.37 | — | — | — | — | — | — | — | |
| DPT (Image)Backbone=DPT, Scaling=Image, Training Dataset=NYUv2, KITTI2024.10 | 14.468 | 0.179 | — | 0.308 | 76.3 | — | — | — | — | — | — | 0.931 | 0.975 | — | |
| AdaBinsEncoder=[54], Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 14.595 | — | 4.791 | — | 78.9 | — | — | — | — | — | — | — | — | — | |
| TC-Depth-FSEncoder=ResNet50, Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 15.121 | — | 5.285 | — | 77.7 | — | — | — | — | — | — | — | — | — | |
| AdaBinsEncoder=[46], Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 15.228 | — | 4.95 | — | 78 | — | — | — | — | — | — | — | — | — | |
| DPT (Global)Backbone=DPT, Scaling=Global, Training Dataset=KITTI2024.10 | 15.967 | 0.183 | — | 0.312 | 75.2 | — | — | — | — | — | — | 0.925 | 0.969 | — | |
| ManyDepth-FSEncoder=ResNet50, Zero-shot evaluation=true, Resolution=1216 x 1936, Evaluation distance (150m)=true2024.03 | 16.123 | — | 5.471 | — | 74.4 | — | — | — | — | — | — | — | — | — | |
| GEDepth-AdaptiveBackbone=DepthFormer2023.09 | 16.132 | 0.261 | — | — | — | — | — | — | — | — | — | — | — | — |