Monocular Depth Estimation on Cityscapes
93.1Accuracy (delta < 1.25)SPIdepth
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| SPIdepthTraining Configuration=-, K→C2024.04 | 93.1 | 98.6 | 99.5 | 0.083 | 0.741 | 5.205 | 0.13 | — | — | |
| Ours-MonoViTsemantic information usage=Training, Resolution=192 x 6402023.12 | 92 | 98.1 | 99.4 | 0.088 | 0.795 | 5.368 | 0.14 | — | — | |
| RM-DepthSemantics=false, Training dataset=CS2023.03 | 91.3 | 98 | 99.3 | 0.09 | 0.825 | 5.503 | 0.143 | — | — | |
| RM-DepthTraining Configuration=MMask, C2024.04 | 91.3 | 98 | 99.3 | 0.09 | 0.825 | 5.503 | 0.143 | — | — | |
| ProDepthTraining Configuration=MMask, C2024.04 | 90.8 | 97.8 | 99.3 | 0.095 | 0.876 | 5.531 | 0.146 | — | — | |
| Ours-CADepthsemantic information usage=Training, Resolution=192 x 6402023.12 | 90.7 | 97.8 | 99.2 | 0.097 | 0.966 | 5.646 | 0.15 | — | — | |
| Ours-MonoViTsemantic information usage=Training, Resolution=128 x 4162023.12 | 90.5 | 97.6 | 99.2 | 0.096 | 0.93 | 5.806 | 0.152 | — | — | |
| Ours-Monodepth2semantic information usage=Training, Resolution=192 x 6402023.12 | 89.6 | 97.3 | 99 | 0.102 | 1.024 | 6.015 | 0.159 | — | — | |
| Ours-HR-Depthsemantic information usage=Training, Resolution=192 x 6402023.12 | 89.6 | 97.4 | 99.1 | 0.1 | 1.01 | 5.998 | 0.157 | — | — | |
| DynamicDepthsemantic information usage=Training and Testing, Resolution=128 x 416, multiple test frames=true2023.12 | 89.5 | 97.4 | 99.1 | 0.103 | 1 | 5.867 | 0.157 | — | — | |
| RM-Depthadditional object motion estimation network=true, Resolution=192 x 6402023.12 | 89.5 | 97.6 | 99.3 | 0.1 | 0.839 | 5.774 | 0.154 | — | — | |
| SQLdepthTraining Configuration=-, K→C2024.04 | 88.8 | 97.2 | 99 | 0.106 | 1.173 | 6.237 | 0.163 | — | — | |
| Ours-Monodepth2semantic information usage=Training, Resolution=128 x 4162023.12 | 88.1 | 96.8 | 98.9 | 0.11 | 1.179 | 6.39 | 0.169 | — | — | |
| MonoViTResolution=192 x 6402023.12 | 88.1 | 97.4 | 99.1 | 0.106 | 1.098 | 6.071 | 0.16 | — | — | |
| JetViT-DepthAnythingSize(Spec)=Giant(2 FA), Latency (ms)=90.65, Throughput (samples/s)=11.39, Zero-shot evaluation protocol=true2026.05 | 87.9 | — | — | 0.109 | — | — | — | — | — | |
| JetViT-DepthAnythingSize(Spec)=Large(2 FA), Latency (ms)=32.63, Throughput (samples/s)=32.13, Zero-shot evaluation protocol=true2026.05 | 87.6 | — | — | 0.108 | — | — | — | — | — | |
| DepthAnythingV2Size(Spec)=Giant, Latency (ms)=164.25, Throughput (samples/s)=6.35, Zero-shot evaluation protocol=true2026.05 | 87.6 | — | — | 0.111 | — | — | — | — | — | |
| ManyDepthSemantic supervision=false, Test frames=2 (-1, 0), WxH=416 x 128, Evaluation crop=A2021.04 | 87.5 | 96.7 | 98.9 | 0.114 | 1.193 | 6.223 | 0.17 | — | — | |
| ManyDepthTraining Configuration=MMask, C2024.04 | 87.5 | 96.7 | 98.9 | 0.114 | 1.193 | 6.223 | 0.17 | — | — | |
| DepthAnythingV2Size(Spec)=Large, Latency (ms)=60.76, Throughput (samples/s)=14.20, Zero-shot evaluation protocol=true2026.05 | 87.2 | — | — | 0.111 | — | — | — | — | — | |
| JetViT-DepthAnythingSize(Spec)=Large(0 FA), Latency (ms)=29.70, Throughput (samples/s)=34.97, Zero-shot evaluation protocol=true2026.05 | 87.2 | — | — | 0.11 | — | — | — | — | — | |
| JetViT-DepthAnythingSize(Spec)=Large(1 FA), Latency (ms)=31.04, Throughput (samples/s)=33.72, Zero-shot evaluation protocol=true2026.05 | 87.1 | — | — | 0.109 | — | — | — | — | — | |
| Lee et al.Semantics=false, Training dataset=CS2023.03 | 86.8 | 96.1 | 98.3 | 0.111 | 1.158 | 6.437 | 0.182 | — | — | |
| InstaDMadditional object motion estimation network=true, semantic information usage=Training, Resolution=256 x 8322023.12 | 86.8 | 96.1 | 98.3 | 0.111 | 1.158 | 6.437 | 0.182 | — | — | |
| InstaDMTraining Configuration=MMask, C2024.04 | 86.8 | 96.1 | 98.3 | 0.111 | 1.158 | 6.437 | 0.182 | — | — | |
| Monodepth2Resolution=192 x 6402023.12 | 86.5 | 96.4 | 98.8 | 0.125 | 1.474 | 6.688 | 0.18 | — | — | |
| CADepthResolution=192 x 6402023.12 | 86.2 | 96.2 | 98.6 | 0.124 | 1.278 | 6.771 | 0.183 | — | — | |
| MonoViTResolution=128 x 4162023.12 | 86 | 96.5 | 99 | 0.114 | 1.238 | 6.589 | 0.174 | — | — | |
| HR-DepthResolution=192 x 6402023.12 | 85.7 | 96.3 | 98.8 | 0.12 | 1.253 | 6.714 | 0.179 | — | — | |
| Lee et al.additional object motion estimation network=true, semantic information usage=Training, Resolution=256 x 8322023.12 | 85.2 | 95.1 | 98.2 | 0.116 | 1.214 | 6.695 | 0.186 | — | — | |
| Lee et al.Training Configuration=MMask, C2024.04 | 85.2 | 95.1 | 98.2 | 0.116 | 1.213 | 6.695 | 0.186 | — | — | |
| Ours (R50)Encoder Backbone=ResNet-502026.05 | 85 | 95.5 | 98.1 | 0.115 | 1.221 | 6.79 | 0.186 | — | — | |
| Monodepth2Semantic supervision=false, Test frames=1, WxH=416 x 128, Evaluation crop=A2021.04 | 84.9 | 95.7 | 98.3 | 0.129 | 1.569 | 6.876 | 0.187 | — | — | |
| Monodepth2Resolution=128 x 4162023.12 | 84.9 | 95.7 | 98.3 | 0.129 | 1.569 | 6.876 | 0.187 | — | — | |
| Monodepth2Training Configuration=-, C2024.04 | 84.9 | 95.7 | 98.3 | 0.129 | 1.569 | 6.876 | 0.187 | — | — | |
| Li et al.Semantic supervision=false, Test frames=1, WxH=416 x 128, Evaluation crop=A2021.04 | 84.6 | 95.2 | 98.2 | 0.119 | 1.29 | 6.98 | 0.19 | — | — | |
| Li et al. [28]Semantics=false, Training dataset=CS2023.03 | 84.6 | 95.2 | 98.2 | 0.119 | 1.29 | 6.98 | 0.19 | — | — | |
| Li et al.additional object motion estimation network=true, Resolution=128 x 4162023.12 | 84.6 | 95.2 | 98.2 | 0.119 | 1.29 | 6.98 | 0.19 | — | — | |
| Li et al.Training Configuration=MMask, C2024.04 | 84.6 | 95.2 | 98.2 | 0.119 | 1.29 | 6.98 | 0.19 | — | — | |
| Li et al.Use of off-the-shelf semantic algorithm=true2026.05 | 84.6 | 95.1 | 98 | 0.119 | 1.29 | 6.98 | 0.19 | — | — | |
| CoopNet2026.05 | 84.6 | 95.1 | 98 | 0.121 | 1.443 | 7.01 | 0.19 | — | — | |
| GLNetSemantics=false, Training dataset=CS, Note=with online refinement2023.03 | 84.3 | 93.8 | 97.6 | 0.129 | 1.044 | 5.361 | 0.212 | — | — | |
| Videos in the WildSemantic supervision=true, Test frames=1, WxH=416 x 128, Evaluation crop=A2021.04 | 83 | 94.7 | 98.1 | 0.127 | 1.33 | 6.96 | 0.195 | — | — | |
| Gordon et al.Semantics=true, Training dataset=CS2023.03 | 83 | 94.7 | 98.1 | 0.127 | 1.33 | 6.96 | 0.195 | — | — | |
| Gordon et al.additional object motion estimation network=true, semantic information usage=Training, Resolution=128 x 4162023.12 | 83 | 94.7 | 98.1 | 0.127 | 1.33 | 6.96 | 0.195 | — | — | |
| Videos in the WildTraining Configuration=MMask, C2024.04 | 83 | 94.7 | 98.1 | 0.127 | 1.33 | 6.96 | 0.195 | — | — | |
| LearnKUse of off-the-shelf semantic algorithm=true2026.05 | 83 | 94.7 | 98.1 | 0.127 | 1.33 | 6.96 | 0.195 | — | — | |
| ManyDepthSemantic supervision=false, Test frames=2 (-1, 0), WxH=512 x 192, Evaluation crop=B2021.04 | 82.7 | 95.4 | 98.5 | 0.137 | 1.578 | 7.249 | 0.197 | — | — | |
| Struct2DepthSemantic supervision=true, Test frames=3 (-1, 0, +1), WxH=416 x 128, Evaluation crop=A2021.04 | 82.6 | 93.7 | 97.2 | 0.151 | 2.492 | 7.024 | 0.202 | — | — | |
| Struct2DepthSemantic supervision=true, Test frames=1, WxH=416 x 128, Evaluation crop=A2021.04 | 81.3 | 94.2 | 97.6 | 0.145 | 1.737 | 7.28 | 0.205 | — | — | |
| Struct2DepthSemantics=true, Training dataset=CS2023.03 | 81.3 | 94.2 | 97.6 | 0.145 | 1.737 | 7.28 | 0.205 | — | — | |
| Struct2Depthadditional object motion estimation network=true, semantic information usage=Training, Resolution=128 x 4162023.12 | 81.3 | 94.2 | 97.6 | 0.145 | 1.737 | 7.28 | 0.205 | — | — | |
| Struct2Depth 2Training Configuration=MMask, C2024.04 | 81.3 | 94.2 | 97.6 | 0.145 | 1.737 | 7.28 | 0.205 | — | — | |
| Struct2DepthUse of off-the-shelf semantic algorithm=true2026.05 | 81.3 | 94.2 | 97.8 | 0.145 | 1.737 | 7.28 | 0.205 | — | — | |
| Monodepth2Semantic supervision=false, Test frames=1, WxH=512 x 192, Evaluation crop=B2021.04 | 78.7 | 94.1 | 97.8 | 0.159 | 2.016 | 8.074 | 0.221 | — | — | |
| Struct2DepthSemantic supervision=false, Test frames=3 (-1, 0, +1), WxH=416 x 128, Evaluation crop=A2021.04 | 77.4 | 90.8 | 95.4 | 0.222 | 5.737 | 8.613 | 0.258 | — | — | |
| MiDaS V3 DPTSize(Spec)=Large(Swin), Latency (ms)=39.28, Throughput (samples/s)=27.14, Zero-shot evaluation protocol=true2026.05 | 74.3 | — | — | 0.186 | — | — | — | — | — | |
| AdaShareBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=21.285, # Params Rel. (%)=-50.02021.10 | 71.4 | 86.8 | 93.1 | 0.018 | — | — | — | 15.5 | 10.9 | |
| Pilzer et al.Semantic supervision=false, Test frames=1, WxH=512 x 2562021.04 | 71 | 87.1 | 93.7 | 0.24 | 4.264 | 8.049 | 0.334 | — | — | |
| MTANBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=51.296, # Params Rel. (%)=+20.52021.10 | 71 | 86.3 | 92.8 | 0.017 | — | — | — | 16.4 | 11.3 | |
| Pilzer et al.Training Configuration=GAN, C2024.04 | 71 | 87.1 | 93.7 | 0.24 | 4.264 | 8.049 | 0.334 | — | — | |
| Cross-StitchBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=42.569, # Params Rel. (%)=+0.02021.10 | 70 | 86.3 | 93.1 | 0.017 | — | — | — | 17.2 | 11.4 | |
| AutoMTLBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=28.819, # Params Rel. (%)=-32.32021.10 | 70 | 86.6 | 93.4 | 0.018 | — | — | — | 17.1 | 13.2 | |
| NDDR-CNNBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=44.059, # Params Rel. (%)=+3.52021.10 | 69.9 | 86.3 | 93 | 0.018 | — | — | — | 15.9 | 11.5 | |
| SluiceBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=42.569, # Params Rel. (%)=+0.02021.10 | 68.9 | 85.8 | 92.8 | 0.018 | — | — | — | 15.3 | 10.1 | |
| DENBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=23.838, # Params Rel. (%)=-44.02021.10 | 68.2 | 84.5 | 91.6 | 0.018 | — | — | — | 11.3 | 6.8 | |
| Multi-TaskBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=21.285, # Params Rel. (%)=-50.02021.10 | 58.8 | 80.5 | 89.9 | 0.026 | — | — | — | 1.5 | 3.1 | |
| Single-TaskBackbone=Deeplab-ResNet-34, Task-specific head architecture=ASPP, # Params Abs. (M)=42.5692021.10 | 57.5 | 76.9 | 87 | 0.026 | — | — | — | — | — | |
| MP@Res5variant=PAG (Pixel-wise Attentional Gating)2018.05 | 34.6 | 66.2 | 77.2 | — | — | — | — | — | — | |
| MultiPoolGating parameter p=0.92018.05 | 34.5 | 65.7 | 76.9 | — | — | — | — | — | — | |
| MP@Res5variant=w-Avg. (Softmax weighted average)2018.05 | 33.7 | 65.9 | 76.9 | — | — | — | — | — | — | |
| MultiPoolGating parameter p=0.72018.05 | 32 | 63.5 | 75.8 | — | — | — | — | — | — | |
| baseline2018.05 | 29 | 53.8 | 75.8 | — | — | — | — | — | — | |
| MultiPoolGating parameter p=0.52018.05 | 28.7 | 58.7 | 71.6 | — | — | — | — | — | — |