Monocular Depth Estimation on KITTI Eigen split
0.048Abs RelScaleDepth-K
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ScaleDepth-Kdepth range=0-80m2024.07 | 0.048 | 0.136 | 1.987 | 0.073 | 98 | 99.8 | 100 | — | — | |
| WorDepthBackbone=Swin-L2024.04 | 0.049 | — | 2.039 | 0.074 | 97.9 | 99.8 | 99.9 | — | — | |
| PixelFormer+MAMoType=VD, Encoder=Swin-large2023.07 | 0.049 | 0.13 | 1.884 | 0.072 | 97.7 | 0.998 | 0.9995 | — | — | |
| Ourscap/m=0–802026.04 | 0.049 | 0.147 | 2.062 | 0.07 | 97.6 | 99.8 | 99.9 | — | — | |
| IEBinsBackbone=Swin-Large, Pre-trained=ImageNet-22K2023.09 | 0.05 | 0.142 | 2.011 | 0.075 | 97.8 | 99.8 | 99.9 | — | — | |
| iDiscdepth range=0-80m2024.07 | 0.05 | 0.145 | 2.067 | 0.077 | 97.7 | 99.7 | 99.9 | — | — | |
| IEBinsdepth range=0-80m2024.07 | 0.05 | 0.142 | 2.011 | 0.075 | 97.8 | 99.8 | 99.9 | — | — | |
| SwinV2-MIMType=SF, Encoder=Swin-large2023.07 | 0.05 | 0.139 | 1.966 | 0.075 | 97.7 | 0.998 | 0.9995 | — | — | |
| URCDCType=SF, Encoder=Swin-large2023.07 | 0.05 | 0.142 | 2.032 | 0.076 | 97.7 | 0.997 | 0.9994 | — | — | |
| NeWCRFs+MAMoType=VD, Encoder=Swin-large2023.07 | 0.05 | 0.141 | 2.003 | 0.076 | 97.7 | 0.998 | 0.9994 | — | — | |
| DDPcap/m=0–802026.04 | 0.05 | 0.148 | — | 0.076 | 97.5 | 99.7 | 99.9 | — | — | |
| PixelFormerBackbone=Swin-Large, Pre-trained=ImageNet-22K2023.09 | 0.051 | 0.149 | 2.081 | 0.077 | 97.6 | 99.7 | 99.9 | — | — | |
| NeWCRFs+MAMoType=VD, Encoder=Swin-Base2023.07 | 0.051 | 0.149 | 2.09 | 0.078 | 97.6 | 0.998 | 0.9994 | — | — | |
| PixelFormerParams=270.9M, MACs=385.4G, Lat.=60.9ms, Thrp.=192026.03 | 0.051 | — | — | — | 97.6 | — | — | — | 0.36 | |
| PixelFormercap/m=0–802026.04 | 0.051 | 0.149 | 2.081 | 0.077 | 97.6 | 99.7 | 99.9 | — | — | |
| NeWCRFsBackbone=Swin-Large, Pre-trained=ImageNet-22K2023.09 | 0.052 | 0.155 | 2.129 | 0.079 | 97.4 | 99.7 | 99.9 | — | — | |
| BinsFormerBackbone=Swin-Large, Pre-trained=ImageNet-22K2023.09 | 0.052 | 0.151 | 2.098 | 0.079 | 97.4 | 99.7 | 99.9 | — | — | |
| NeWCRFsBackbone=Swin-L2024.04 | 0.052 | — | 2.129 | 0.079 | 97.4 | 99.7 | 99.9 | — | — | |
| DepthFormerBackbone=Swin-L2024.04 | 0.052 | — | 2.143 | 0.079 | 97.5 | 99.7 | 99.9 | — | — | |
| NeWCRFsdepth range=0-80m2024.07 | 0.052 | 0.155 | 2.129 | 0.079 | 97.4 | 99.7 | 99.9 | — | — | |
| BinsFormerdepth range=0-80m2024.07 | 0.052 | 0.151 | 2.098 | 0.079 | 97.4 | 99.7 | 99.9 | — | — | |
| DPTScaling=Linear fit, Training Dataset=KITTI2024.10 | 0.052 | — | 2.198 | 0.08 | 97.4 | 99.7 | 99.9 | — | — | |
| BinsFormerType=SF, Encoder=Swin-large2023.07 | 0.052 | 0.151 | 2.098 | 0.079 | 97.5 | 0.997 | 0.9992 | — | — | |
| PixelFormerType=VD, Encoder=Swin-large2023.07 | 0.052 | 0.152 | 2.093 | 0.079 | 97.5 | 0.997 | 0.9994 | — | — | |
| DepthFormer2026.01 | 0.052 | 0.158 | 2.143 | 0.079 | 97.5 | 99.7 | 99.9 | — | — | |
| NeWCRFs2026.01 | 0.052 | 0.155 | 2.129 | 0.079 | 97.4 | 99.7 | 99.9 | — | — | |
| DPT-HybridScaling=Linear Fit, Train Dataset=KITTI2026.01 | 0.052 | — | 2.198 | 0.08 | 97.4 | 99.7 | 99.9 | — | — | |
| BinsFormerParams=254.6M, MACs=409.4G, Lat.=97.5ms, Thrp.=132026.03 | 0.052 | — | — | — | 97.4 | — | — | — | 0.38 | |
| NeWCRFsParams=270.4M, MACs=395.9G, Lat.=74.8ms, Thrp.=142026.03 | 0.052 | — | — | — | 97.5 | — | — | — | 0.36 | |
| BinsFormercap/m=0–802026.04 | 0.052 | 0.151 | 2.098 | 0.079 | 97.5 | — | — | — | — | |
| NeWCRFscap/m=0–802026.04 | 0.052 | 0.155 | 2.129 | 0.079 | 97.4 | 99.7 | 99.9 | — | — | |
| NeWCRFsType=VD, Encoder=Swin-large2023.07 | 0.053 | 0.154 | 2.118 | 0.08 | 97.4 | 0.997 | 0.9994 | — | — | |
| Yu et al.Backbone=Swin-L2024.04 | 0.054 | — | 2.134 | 0.081 | 97.2 | 99.6 | 99.9 | — | — | |
| BaselineBackbone=Swin-L2024.04 | 0.054 | — | 2.343 | 0.085 | 96.9 | 99.6 | 99.9 | — | — | |
| ZoeDepthScaling=Image, Training Dataset=KITTI2024.10 | 0.054 | — | 2.281 | 0.082 | 97.1 | 99.6 | 99.9 | — | — | |
| NeWCRFsType=VD, Encoder=Swin-Base2023.07 | 0.054 | 0.157 | 2.14 | 0.081 | 97.3 | 0.997 | 0.9993 | — | — | |
| ZoeDepthScaling=Image, Train Dataset=KITTI2026.01 | 0.054 | — | 2.281 | 0.082 | 97.1 | 99.6 | 99.9 | — | — | |
| EvoNet-K3Params=26.3M, MACs=45.0G, Lat.=28.0ms, Thrp.=652026.03 | 0.054 | — | — | — | 96.9 | — | — | — | 3.68 | |
| ZoeDepthcap/m=0–802026.04 | 0.054 | 0.189 | 2.44 | 0.083 | 97 | 99.6 | 99.9 | — | — | |
| IEBinsBackbone=Swin-Tiny2023.09 | 0.056 | 0.169 | 2.205 | 0.084 | 97 | 99.6 | 99.9 | — | — | |
| IEBinsParams=90.7M, MACs=527.3G, Lat.=69.5ms, Thrp.=162026.03 | 0.056 | — | — | — | 97 | — | — | — | 1.07 | |
| EvoNet-K2Params=22.6M, MACs=36.2G, Lat.=24.6ms, Thrp.=832026.03 | 0.056 | — | — | — | 96.6 | — | — | — | 4.28 | |
| CaBinsText Encoder (at Inference)=true, Vision Encoder (Finetuned)=true2026.01 | 0.057 | 0.186 | 2.322 | 0.088 | 96.4 | 99.5 | 99.9 | — | — | |
| AdaBinsBackbone=EfficientNet-B5+mini-ViT2023.09 | 0.058 | 0.19 | 2.36 | 0.088 | 96.4 | 99.5 | 99.9 | — | — | |
| BinsFormerBackbone=Swin-Tiny2023.09 | 0.058 | 0.183 | 2.286 | 0.088 | 96.8 | 99.5 | 99.9 | — | — | |
| ASTransformerBackbone=ViT-B2024.04 | 0.058 | — | 2.685 | 0.089 | 96.3 | 99.5 | 99.9 | — | — | |
| AdaBinsBackbone=EffNet-B5+ViT-mini2024.04 | 0.058 | — | 2.36 | 0.089 | 96.4 | 99.5 | 99.9 | — | — | |
| AdaBinsType=SF, Encoder=EfficientNet-B5+mViT2023.07 | 0.058 | 0.19 | 2.36 | 0.088 | 96.4 | 0.995 | 0.9991 | — | — | |
| DepthFormerType=SF, Encoder=MiT-B42023.07 | 0.058 | 0.187 | 2.285 | 0.087 | 96.7 | 0.996 | 0.9991 | — | — | |
| ASTransformer2026.01 | 0.058 | — | 2.685 | 0.089 | 96.3 | 99.5 | 99.9 | — | — | |
| iDiscParams=40.7M, MACs=256.2G, Lat.=83.4ms, Thrp.=142026.03 | 0.058 | — | — | — | 96.8 | — | — | — | 2.38 | |
| AdaBinscap/m=0–802026.04 | 0.058 | 0.19 | 2.36 | 0.088 | 96.4 | 99.5 | 99.9 | — | — | |
| Language-guided Depth Calibration Framework (DPT-Hybrid)Scaling=Ours, Train Dataset=NYUv2, KITTI2026.01 | 0.059 | — | 2.32 | 0.088 | 96.4 | 99.5 | 99.9 | — | — | |
| BTScap/m=0–802026.04 | 0.059 | 0.241 | 2.756 | 0.09 | 95.6 | 99.3 | 99.8 | — | — | |
| BTSBackbone=DenseNet-1612023.09 | 0.06 | 0.249 | 2.798 | 0.096 | 95.5 | 99.3 | 99.8 | — | — | |
| PWABackbone=ResNeXt-1012023.09 | 0.06 | 0.221 | 2.604 | 0.093 | 95.8 | 99.4 | 99.9 | — | — | |
| Big to SmallBackbone=DenseNet-1612024.04 | 0.06 | — | 2.798 | 0.096 | 95.5 | 99.3 | 99.8 | — | — | |
| BTSdepth range=0-80m2024.07 | 0.06 | 0.249 | 2.798 | 0.096 | 95.5 | 99.3 | 99.8 | — | — | |
| AdaBinsdepth range=0-80m2024.07 | 0.06 | 0.197 | 2.372 | 0.09 | 96.3 | 99.5 | 99.9 | — | — | |
| RSA (DPT)Scaling=RSA (Ours), Backbone=DPT, Training Dataset=NYUv2, KITTI2024.10 | 0.06 | — | 2.342 | 0.089 | 96.2 | 99.4 | 99.8 | — | — | |
| ManyDepth-FSType=MF, Encoder=Swin-large, Supervision=fully-supervised2023.07 | 0.06 | 0.248 | 2.747 | 0.099 | 95.5 | 0.993 | 0.9981 | — | — | |
| Language-guided Depth Calibration Framework (DPT-Hybrid)Scaling=Ours, Train Dataset=KITTI2026.01 | 0.06 | — | 2.33 | 0.089 | 96.5 | 99.5 | 99.9 | — | — | |
| DPT-HybridScaling=RSA, Train Dataset=NYUv2, KITTI2026.01 | 0.06 | — | 2.342 | 0.089 | 96.2 | 99.4 | 99.8 | — | — | |
| EvoNet-K1Params=18.0M, MACs=27.3G, Lat.=18.6ms, Thrp.=1172026.03 | 0.06 | — | — | — | 96 | — | — | — | 5.34 | |
| PWAcap/m=0–802026.04 | 0.06 | 0.221 | 2.604 | 0.093 | 95.8 | 99.4 | 99.9 | — | — | |
| RSA (DPT)Scaling=RSA (Ours), Backbone=DPT, Training Dataset=KITTI2024.10 | 0.061 | — | 2.354 | 0.09 | 96.3 | 99.5 | 99.9 | — | — | |
| DPT-HybridScaling=RSA, Train Dataset=KITTI2026.01 | 0.061 | — | 2.354 | 0.09 | 96.3 | 99.5 | 99.9 | — | — | |
| BTSParams=66.5M, MACs=184.3G, Lat.=32.9ms, Thrp.=352026.03 | 0.061 | — | — | — | 95.4 | — | — | — | 1.43 | |
| DPT-HybirdBackbone=ViT-B2024.04 | 0.062 | — | 2.573 | 0.092 | 95.9 | 99.5 | 99.9 | — | — | |
| DPTScaling=Global, Training Dataset=KITTI2024.10 | 0.062 | — | 2.575 | 0.092 | 95.9 | 99.5 | 99.9 | — | — | |
| DPTParams=123.1M, MACs=319.3G, Lat.=60.1ms, Thrp.=192026.03 | 0.062 | — | — | — | 95.9 | — | — | — | 0.78 | |
| DPT*cap/m=0–802026.04 | 0.062 | — | 2.573 | 0.092 | 95.9 | 99.5 | 99.9 | — | — | |
| TransDepthBackbone=ResNet-50+ViT-B/16+2023.09 | 0.064 | 0.252 | 2.755 | 0.098 | 95.6 | 99.4 | 99.9 | — | — | |
| TransDepthBackbone=ViT-B2024.04 | 0.064 | — | 2.755 | 0.098 | 95.6 | 99.4 | 99.9 | — | — | |
| DPTScaling=Image, Training Dataset=KITTI2024.10 | 0.064 | — | 2.379 | 0.092 | 96.1 | 99.5 | 99.9 | — | — | |
| RSA (DPT)Scaling=RSA (Ours), Backbone=DPT, Training Dataset=NYUv2, KITTI, VOID2024.10 | 0.064 | — | 2.335 | 0.091 | 96.1 | 99.4 | 99.9 | — | — | |
| DPT-HybridScaling=Image, Train Dataset=KITTI2026.01 | 0.064 | — | 2.379 | 0.092 | 96.1 | 99.5 | 99.9 | — | — | |
| DPTScaling=Image, Training Dataset=NYUv2, KITTI2024.10 | 0.066 | — | 2.477 | 0.098 | 95.6 | 98.9 | 99.3 | — | — | |
| DPT-HybridScaling=Image, Train Dataset=NYUv2, KITTI2026.01 | 0.066 | — | 2.477 | 0.098 | 95.6 | 98.9 | 99.3 | — | — | |
| AdaBinsParams=78.3M, MACs=261.7G, Lat.=47.2ms, Thrp.=212026.03 | 0.067 | — | — | — | 94.9 | — | — | — | 1.21 | |
| DPTScaling=Image, Training Dataset=NYUv2, KITTI, VOID2024.10 | 0.068 | — | 2.568 | 0.098 | 95.2 | 98.7 | 99.3 | — | — | |
| DPTScaling=Median, Training Dataset=KITTI2024.10 | 0.069 | — | 3.365 | 0.1 | 95 | 99.4 | 99.9 | — | — | |
| ManyDepth-FSType=MF, Encoder=ResNet50, Supervision=fully-supervised2023.07 | 0.069 | 0.342 | 3.414 | 0.111 | 93 | 0.989 | 0.997 | — | — | |
| P3DepthBackbone=ResNet-1012023.09 | 0.071 | 0.27 | 2.842 | 0.103 | 95.3 | 99.3 | 99.8 | — | — | |
| TC-Depth-FSType=MF, Encoder=ResNet50, Supervision=fully-supervised2023.07 | 0.071 | 0.33 | 3.222 | 0.108 | 92.2 | 0.993 | 0.997 | — | — | |
| ResNet-DPT+MAMoType=VD, Encoder=ResNet502023.07 | 0.071 | 0.301 | 2.984 | 0.121 | 92.6 | 0.99 | 0.9971 | — | — | |
| P3DepthParams=94.3M, MACs=401.8G, Lat.=36.1ms, Thrp.=342026.03 | 0.071 | — | — | — | 95.3 | — | — | — | 1.01 | |
| DORNBackbone=ResNet-1012023.09 | 0.072 | 0.307 | 2.727 | 0.12 | 93.2 | 98.4 | 99.4 | — | — | |
| VNLBackbone=ResNeXt-1012023.09 | 0.072 | — | 3.258 | 0.117 | 93.8 | 99 | 99.8 | — | — | |
| DORNBackbone=ResNet-1012024.04 | 0.072 | — | 2.727 | 0.12 | 93.2 | 98.4 | 99.5 | — | — | |
| Yin et al.Backbone=ResNeXt-1012024.04 | 0.072 | — | 3.258 | 0.117 | 93.8 | 99 | 99.8 | — | — | |
| DORNdepth range=0-80m2024.07 | 0.072 | 0.307 | 2.727 | 0.12 | 93.2 | 98.4 | 99.4 | — | — | |
| DORN2026.01 | 0.072 | 0.307 | 2.727 | 0.12 | 93.2 | 98.4 | 99.4 | — | — | |
| CLIP2DepthText Encoder (at Inference)=true, Vision Encoder (Finetuned)=false2026.01 | 0.074 | 0.303 | 2.948 | — | 93.8 | 99 | 99.8 | — | — | |
| Depth Anythingzero-shot=true, Num. Train Samples=63.5M, Architecture Type=Transformer2024.09 | 0.076 | — | — | — | 94.7 | — | — | 2.2 | — | |
| PrimeDepthzero-shot=true, Num. Train Samples=74K, Architecture Type=Single-step diffusion2024.09 | 0.079 | — | — | — | 93.7 | — | — | 2 | — | |
| Flow2DepthType=MF, Encoder=CNN, multiple networks=true2023.07 | 0.081 | 0.488 | 3.651 | 0.146 | 91.2 | 0.97 | 0.9883 | — | — | |
| ResNet-DPTType=VD, Encoder=ResNet502023.07 | 0.085 | 0.383 | 3.242 | 0.13 | 91.3 | 0.981 | 0.996 | — | — | |
| ProDepthTest frames=2 (-1, 0), Semantics=false, W x H=640 x 1922024.07 | 0.086 | 0.629 | 4.139 | 0.166 | 91.8 | 96.9 | 98.4 | — | — | |
| DepthFormerTest frames=2 (-1, 0), Semantics=false, W x H=640 x 1922024.07 | 0.09 | 0.661 | 4.149 | 0.175 | 90.5 | 96.7 | 98.4 | — | — |