Depth Prediction on Cityscapes (test)
5.443RMSEPilzer et al. (Row 1)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Pilzer et al. (Row 1)Train=Cityscapes, Range cap=80m, Cropping=Unspecified/Different cropping2019.06 | 5.443 | 0.44 | 6.036 | 0.398 | 0.73 | 0.887 | 0.944 | |
| SwinMTLResolution=512x1024, Supervision=Supervised2024.03 | 5.481 | 0.089 | 1.051 | 0.139 | 0.921 | 0.976 | 0.99 | |
| Pilzer et al. (Row 2)Train=Cityscapes, Range cap=80m, Cropping=Unspecified/Different cropping2019.06 | 5.741 | 0.467 | 7.399 | 0.493 | 0.735 | 0.89 | 0.945 | |
| Pilzer et al.Resolution=512x256, Supervision=Self-supervised2024.03 | 5.745 | 0.438 | 5.713 | 0.4 | 0.711 | 0.877 | 0.94 | |
| DynamicDepthTest frames=2 (-1, 0), WxH=416 x 1282022.03 | 5.867 | 0.103 | 1 | 0.157 | 0.895 | 0.974 | 0.991 | |
| DynamicDepthResolution=416x128, Supervision=Self-supervised2024.03 | 5.867 | 0.103 | 1 | 0.157 | 0.895 | 0.974 | 0.991 | |
| ManyDepthTest frames=2 (-1, 0), Semantics=No, WxH=416 x 1282021.04 | 6.223 | 0.114 | 1.193 | 0.17 | 0.875 | 0.967 | 0.989 | |
| ManyDepthTest frames=2 (-1, 0), WxH=416 x 1282022.03 | 6.223 | 0.114 | 1.193 | 0.17 | 0.875 | 0.967 | 0.989 | |
| SwinMTLResolution=1024x2048, Backbone=SwinV2-B2024.03 | 6.352 | — | — | — | — | — | — | |
| Insta-DMBackbone=ResNet18, Training=C(S), Semantic Knowledge=true2021.02 | 6.437 | 0.111 | 1.158 | 0.182 | 0.868 | 0.961 | 0.983 | |
| InstaDMTest frames=1, WxH=832 x 2562022.03 | 6.437 | 0.111 | 1.158 | 0.182 | 0.868 | 0.961 | 0.983 | |
| GLNetResolution=1024x2048, Supervision=Supervised2024.03 | 6.437 | 0.111 | — | 0.182 | 0.868 | 0.961 | 0.983 | |
| InstaDMResolution=832x256, Supervision=Self-supervised2024.03 | 6.437 | 0.111 | 1.158 | 0.182 | 0.868 | 0.961 | 0.983 | |
| 3-waysResolution=1024x20482024.03 | 6.528 | — | — | — | — | — | — | |
| [44]Resolution=1024x20482024.03 | 6.649 | — | — | — | — | — | — | |
| PanopticDepthutilize panoptic segmentation annotations=true2022.06 | 6.69 | — | — | — | — | — | — | |
| Lee et al.Test frames=1, WxH=832 x 2562022.03 | 6.695 | 0.116 | 1.213 | 0.186 | 0.852 | 0.951 | 0.982 | |
| Lee et al.Resolution=832x256, Supervision=Self-supervised2024.03 | 6.695 | 0.116 | 1.213 | 0.186 | 0.852 | 0.951 | 0.982 | |
| PAD-NetResolution=1024x20482024.03 | 6.777 | — | — | — | — | — | — | |
| MTLResolution=1024x20482024.03 | 6.797 | — | — | — | — | — | — | |
| Monodepth2Test frames=1, Semantics=No, WxH=416 x 1282021.04 | 6.876 | 0.129 | 1.569 | 0.187 | 0.849 | 0.957 | 0.983 | |
| Monodepth2Test frames=1, WxH=416 x 1282022.03 | 6.876 | 0.129 | 1.569 | 0.187 | 0.849 | 0.957 | 0.983 | |
| CFCNetInput=RGB+SD, Training Dataset=Cityscapes, Sparse Points=100 pts, Depth Evaluation Cap=50m2019.06 | 6.887 | — | — | — | 0.889 | 0.961 | 0.981 | |
| SDC-Depthutilize panoptic segmentation annotations=true2022.06 | 6.92 | — | — | — | — | — | — | |
| Gordon et al.Backbone=ResNet18, Training=C(S), Semantic Knowledge=true2021.02 | 6.96 | 0.127 | 1.33 | 0.195 | 0.83 | 0.947 | 0.981 | |
| Videos in the WildTest frames=1, Semantics=No, WxH=416 x 1282021.04 | 6.96 | 0.127 | 1.33 | 0.195 | 0.83 | 0.947 | 0.981 | |
| Videos in the WildTest frames=1, WxH=416 x 1282022.03 | 6.96 | 0.127 | 1.33 | 0.195 | 0.83 | 0.947 | 0.981 | |
| Li et al.Backbone=ResNet18, Training=C, Semantic Knowledge=false2021.02 | 6.98 | 0.119 | 1.29 | 0.19 | 0.846 | 0.952 | 0.982 | |
| Li et al.Test frames=1, Semantics=No, WxH=416 x 1282021.04 | 6.98 | 0.119 | 1.29 | 0.19 | 0.846 | 0.952 | 0.982 | |
| Li et al.Test frames=1, WxH=416 x 1282022.03 | 6.98 | 0.119 | 1.29 | 0.19 | 0.846 | 0.952 | 0.982 | |
| Li et al.Resolution=416x128, Supervision=Self-supervised2024.03 | 6.98 | 0.119 | 1.29 | 0.19 | 0.846 | 0.952 | 0.982 | |
| Struct2Depth (M+R)Train=Cityscapes, Range cap=80m, Mode=Motion model + Online refinement2019.06 | 7.0237 | 0.1511 | 2.4916 | 0.2023 | 0.8255 | 0.9372 | 0.9721 | |
| Struct2Depth 2Test frames=3 (-1, 0, +1), Semantics=Yes, WxH=416 x 1282021.04 | 7.024 | 0.151 | 2.492 | 0.202 | 0.826 | 0.937 | 0.972 | |
| Struct2Depth 2Test frames=3 (-1, 0, +1), WxH=416 x 1282022.03 | 7.024 | 0.151 | 2.492 | 0.202 | 0.826 | 0.937 | 0.972 | |
| Zhang et al.2022.06 | 7.1 | — | — | — | — | — | — | |
| Pad-Net2022.06 | 7.12 | — | — | — | — | — | — | |
| Laina et al.2022.06 | 7.27 | — | — | — | — | — | — | |
| Struct2Depth (M)Train=Cityscapes, Range cap=80m, Mode=Motion model2019.06 | 7.2798 | 0.1454 | 1.7368 | 0.2046 | 0.813 | 0.9415 | 0.9775 | |
| Struct2DepthBackbone=ResNet18, Training=C(S), Semantic Knowledge=true2021.02 | 7.28 | 0.145 | 1.737 | 0.205 | 0.813 | 0.942 | 0.978 | |
| Struct2Depth 2Test frames=1, Semantics=Yes, WxH=416 x 1282021.04 | 7.28 | 0.145 | 1.737 | 0.205 | 0.813 | 0.942 | 0.976 | |
| Struct2Depth 2Test frames=1, WxH=416 x 1282022.03 | 7.28 | 0.145 | 1.737 | 0.205 | 0.813 | 0.942 | 0.976 | |
| Struct2DepthResolution=416x128, Supervision=Self-supervised2024.03 | 7.28 | 0.145 | 1.737 | 0.205 | 0.813 | 0.942 | 0.976 | |
| Pilzer et al.Test frames=1, Semantics=No, WxH=512 x 2562021.04 | 8.049 | 0.24 | 4.264 | 0.334 | 0.71 | 0.871 | 0.937 | |
| Pilzer et al.Test frames=1, WxH=512 x 2562022.03 | 8.049 | 0.24 | 4.264 | 0.334 | 0.71 | 0.871 | 0.937 | |
| MGMNetResolution=1024x20482024.03 | 8.3 | — | — | — | — | — | — | |
| Struct2Depth 2Test frames=3 (-1, 0, +1), Semantics=No, WxH=416 x 1282021.04 | 8.613 | 0.222 | 5.737 | 0.258 | 0.774 | 0.908 | 0.954 | |
| Struct2Depth (R)Train=Cityscapes, Range cap=80m, Mode=Online refinement model2019.06 | 8.6133 | 0.2218 | 5.7374 | 0.2584 | 0.7738 | 0.9076 | 0.9542 | |
| CFCNetInput=RGB+SD, Training Dataset=Cityscapes, Sparse Points=50 pts, Depth Evaluation Cap=50m2019.06 | 9.019 | — | — | — | 0.828 | 0.941 | 0.972 | |
| S2R-DepthNetSource Dataset=vKITTI, Cross-dataset=true2021.04 | 11.164 | 0.208 | 2.944 | 0.314 | 0.663 | 0.86 | 0.941 | |
| HybridnetResolution=1024x20482024.03 | 12.09 | — | — | — | — | — | — | |
| T2NetSource Dataset=vKITTI, Cross-dataset=true2021.04 | 13.922 | 0.294 | 4.639 | 0.425 | 0.528 | 0.76 | 0.868 | |
| BaselineSource Dataset=vKITTI, Cross-dataset=true2021.04 | 13.938 | 0.297 | 5.077 | 0.452 | 0.516 | 0.747 | 0.854 | |
| Adaptive Dual-Path FrameworkλCTS=12026.05 | — | 0.0271 | — | — | 0.6269 | — | 0.899 | |
| Standard SemCom2026.05 | — | 0.0213 | — | — | 0.7071 | — | 0.9349 |