Monocular Depth Estimation on KITTI 2015 (Eigen split)
0.072Abs RelDORN
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| DORNPost-processing=false, Data=D, Resolution=385 × 513 crop2019.09 | 0.072 | 0.307 | 2.727 | 0.12 | 93.2 | 98.4 | 99.4 | |
| SVSM finetunedPost-processing=false, Data=D*DS, Resolution=192 × 640 crop2019.09 | 0.094 | 0.626 | 4.252 | 0.177 | 89.1 | 96.5 | 98.4 | |
| Ours HR Resnet50Post-processing=true, Data=S, Resolution=320 × 10242019.09 | 0.096 | 0.71 | 4.393 | 0.185 | 89 | 96.2 | 98.1 | |
| CADepth-NetTrain=MS, Resolution=1024 × 3202021.12 | 0.096 | 0.694 | 4.264 | 0.173 | 90.8 | 96.8 | 98.4 | |
| DVSOPost-processing=true, Data=D+S, Resolution=256 × 5122019.09 | 0.097 | 0.734 | 4.442 | 0.187 | 88.8 | 95.8 | 98 | |
| Guo StereoSupFTAll -> Mono ptPost-processing=false, Data=D*DS, Resolution=256 × 5122019.09 | 0.097 | 0.653 | 4.17 | 0.17 | 88.9 | 96.7 | 98.6 | |
| Ours HRPost-processing=true, Data=MS, Resolution=320 × 10242019.09 | 0.098 | 0.702 | 4.398 | 0.183 | 88.7 | 96.3 | 98.3 | |
| DepthHintsTrain=MS, Resolution=1024 × 3202021.12 | 0.098 | 0.702 | 4.398 | 0.183 | 88.7 | 96.3 | 98.3 | |
| Guo StereoUnsupFT -> Mono ptPost-processing=false, Data=D*S, Resolution=256 × 5122019.09 | 0.099 | 0.745 | 4.424 | 0.182 | 88.4 | 96.3 | 98.3 | |
| Ours HRPost-processing=true, Data=S, Resolution=320 × 10242019.09 | 0.099 | 0.723 | 4.445 | 0.187 | 88.6 | 96.2 | 98.1 | |
| FeatDepthTrain=MS, Resolution=1024 × 3202021.12 | 0.099 | 0.697 | 4.427 | 0.184 | 88.9 | 96.3 | 98.2 | |
| HR-DepthTrain=MS, Resolution=1024 × 3202021.12 | 0.101 | 0.716 | 4.395 | 0.179 | 89.9 | 96.6 | 98.3 | |
| SVSM w/o finetuningPost-processing=false, Data=D*S, Resolution=192 × 640 crop2019.09 | 0.102 | 0.7 | 4.681 | 0.2 | 87.2 | 95.4 | 97.8 | |
| Ours Resnet50Post-processing=true, Data=S, Resolution=192 × 6402019.09 | 0.102 | 0.762 | 4.602 | 0.189 | 88 | 96 | 98.1 | |
| CADepth-NetTrain=M, Resolution=1024 × 3202021.12 | 0.102 | 0.734 | 4.407 | 0.178 | 89.8 | 96.6 | 98.4 | |
| CADepth-NetTrain=M, Resolution=1280 × 3842021.12 | 0.102 | 0.715 | 4.312 | 0.176 | 90 | 96.8 | 98.4 | |
| CADepth-NetTrain=MS, Resolution=640 × 1922021.12 | 0.102 | 0.752 | 4.504 | 0.181 | 89.4 | 96.4 | 98.3 | |
| Monodepth2Post-processing=true, Data=MS, Resolution=192 × 6402019.09 | 0.104 | 0.786 | 4.687 | 0.194 | 87.6 | 95.8 | 98 | |
| Monodepth2Post-processing=true, Data=MS, Resolution=320 × 10242019.09 | 0.104 | 0.775 | 4.562 | 0.191 | 87.8 | 95.9 | 98.1 | |
| FeatDepthTrain=M, Resolution=1024 × 3202021.12 | 0.104 | 0.729 | 4.481 | 0.179 | 89.3 | 96.5 | 98.4 | |
| HR-DepthTrain=M, Resolution=1280 × 3842021.12 | 0.104 | 0.727 | 4.41 | 0.179 | 89.4 | 96.6 | 98.4 | |
| Monodepth2Post-processing=true, Data=S, Resolution=320 × 10242019.09 | 0.105 | 0.822 | 4.692 | 0.199 | 87.6 | 95.4 | 97.7 | |
| OursPost-processing=true, Data=MS, Resolution=192 × 6402019.09 | 0.105 | 0.769 | 4.627 | 0.189 | 87.5 | 95.9 | 98.2 | |
| CADepth-NetTrain=M, Resolution=640 × 1922021.12 | 0.105 | 0.769 | 4.535 | 0.181 | 89.2 | 96.4 | 98.3 | |
| DepthHintsTrain=MS, Resolution=640 × 1922021.12 | 0.105 | 0.769 | 4.627 | 0.189 | 87.5 | 95.9 | 98.2 | |
| OursPost-processing=true, Data=S, Resolution=192 × 6402019.09 | 0.106 | 0.78 | 4.695 | 0.193 | 87.5 | 95.8 | 98 | |
| Johnston et al.Train=M, Resolution=640 × 1922021.12 | 0.106 | 0.861 | 4.699 | 0.185 | 88.9 | 96.2 | 98.2 | |
| HR-DepthTrain=M, Resolution=1024 × 3202021.12 | 0.106 | 0.755 | 4.472 | 0.181 | 89.2 | 96.6 | 98.4 | |
| Monodepth2Train=MS, Resolution=640 × 1922021.12 | 0.106 | 0.818 | 4.75 | 0.196 | 87.4 | 95.7 | 97.9 | |
| Monodepth2Train=MS, Resolution=1024 × 3202021.12 | 0.106 | 0.806 | 4.63 | 0.193 | 87.6 | 95.8 | 98 | |
| DVSO SimpleNetPost-processing=true, Data=DS, Resolution=256 × 5122019.09 | 0.107 | 0.852 | 4.785 | 0.199 | 86.6 | 95 | 97.8 | |
| PackNet-SfMTrain=M, Resolution=1280 × 3842021.12 | 0.107 | 0.802 | 4.538 | 0.186 | 88.9 | 96.2 | 98.1 | |
| HR-DepthTrain=MS, Resolution=640 × 1922021.12 | 0.107 | 0.785 | 4.612 | 0.185 | 88.7 | 96.2 | 98.2 | |
| Kuznietsov et al. [20]Dataset=K(D+L+R), cap=1-50m2018.08 | 0.108 | 0.595 | 3.518 | 0.179 | 87.5 | 96.4 | 98.8 | |
| Monodepth2Post-processing=true, Data=S, Resolution=192 × 6402019.09 | 0.108 | 0.842 | 4.891 | 0.207 | 86.6 | 94.9 | 97.6 | |
| HR-DepthTrain=M, Resolution=640 × 1922021.12 | 0.109 | 0.792 | 4.632 | 0.185 | 88.4 | 96.2 | 98.3 | |
| monoResMatchPost-processing=true, Data=S, Resolution=256 × 512 crop2019.09 | 0.111 | 0.867 | 4.714 | 0.199 | 86.4 | 95.4 | 97.9 | |
| PackNet-SfMTrain=M, Resolution=640 × 1922021.12 | 0.111 | 0.785 | 4.601 | 0.189 | 87.8 | 96 | 98.2 | |
| SuperDepthPost-processing=false, Data=S, Resolution=384 × 10242019.09 | 0.112 | 0.875 | 4.958 | 0.207 | 85.2 | 94.7 | 97.7 | |
| Ours HR Resnet50 w/o pretrainingPost-processing=true, Data=S, Resolution=320 × 10242019.09 | 0.112 | 0.857 | 4.807 | 0.203 | 86.1 | 95.2 | 97.8 | |
| KuznietsovPost-processing=false, Data=DS, Resolution=187 × 6212019.09 | 0.113 | 0.741 | 4.621 | 0.189 | 86.2 | 96 | 98.6 | |
| TrianFlowTrain=M, Resolution=832 × 2562021.12 | 0.113 | 0.704 | 4.581 | 0.184 | 87.1 | 96.1 | 98.4 | |
| SGDepthTrain=M, Resolution=1280 × 3842021.12 | 0.113 | 0.88 | 4.695 | 0.192 | 88.4 | 96.1 | 98.1 | |
| Our fr, all-realDataset=K(I+D), cap=1-50m2018.08 | 0.114 | 0.627 | 3.549 | 0.178 | 86.7 | 96 | 98.6 | |
| Monodepth2Train=M, Resolution=640 × 1922021.12 | 0.115 | 0.903 | 4.863 | 0.193 | 87.7 | 95.9 | 98.1 | |
| Monodepth2Train=M, Resolution=1024 × 3202021.12 | 0.115 | 0.882 | 4.701 | 0.19 | 87.9 | 96.1 | 98.2 | |
| CADepth-NetTrain=M, Resolution=416 × 1282021.12 | 0.116 | 0.893 | 4.906 | 0.192 | 87.4 | 95.7 | 98.1 | |
| Godard et al. [10]Dataset=CS+K(L+R), cap=1-50m2018.08 | 0.117 | 0.762 | 3.972 | 0.206 | 86 | 94.8 | 97.6 | |
| SGDepthTrain=M, Resolution=640 × 1922021.12 | 0.117 | 0.907 | 4.844 | 0.196 | 87.5 | 95.8 | 98 | |
| Ours Resnet50 w/o pretrainingPost-processing=true, Data=S, Resolution=192 × 6402019.09 | 0.118 | 0.941 | 5.055 | 0.21 | 85 | 94.8 | 97.6 | |
| DualNetTrain=M, Resolution=1248 × 3842021.12 | 0.121 | 0.837 | 4.945 | 0.197 | 85.3 | 95.5 | 98.2 | |
| Godard et al.Supervised=Pose, Dataset=CS + K2018.03 | 0.124 | 1.076 | 5.311 | 0.219 | 84.7 | 94.2 | 97.3 | |
| 3Net (Resnet50)Post-processing=true, Data=S, Resolution=256 × 5122019.09 | 0.126 | 0.961 | 5.205 | 0.22 | 83.5 | 94.1 | 97.4 | |
| StrATPost-processing=false, Data=S, Resolution=256 × 5122019.09 | 0.128 | 1.019 | 5.403 | 0.227 | 82.7 | 93.5 | 97.1 | |
| Monodepth2 (w/o pretraining)Post-processing=true, Data=S, Resolution=192 × 6402019.09 | 0.128 | 1.089 | 5.385 | 0.229 | 83.2 | 93.4 | 96.9 | |
| EPC++Post-processing=false, Data=MS, Resolution=256 × 8322019.09 | 0.128 | 0.935 | 5.011 | 0.209 | 83.1 | 94.5 | 97.9 | |
| SGDepthTrain=M, Resolution=416 × 1282021.12 | 0.128 | 1.003 | 5.085 | 0.206 | 85.3 | 95.1 | 97.8 | |
| Monodepth2Train=M, Resolution=416 × 1282021.12 | 0.128 | 1.087 | 5.171 | 0.204 | 85.5 | 95.3 | 97.8 | |
| EPC++Train=MS, Resolution=832 × 2562021.12 | 0.128 | 0.935 | 5.011 | 0.209 | 83.1 | 94.5 | 97.9 | |
| SIGNetSupervised=No2018.12 | 0.133 | 0.905 | 5.181 | 0.208 | 82.5 | 94.7 | 98.1 | |
| SIGNetTrain=M, Resolution=416 × 1282021.12 | 0.133 | 0.905 | 5.181 | 0.208 | 82.5 | 94.7 | 98.1 | |
| ZhanPost-processing=false, Data=MS, Resolution=160 × 6082019.09 | 0.135 | 1.132 | 5.585 | 0.229 | 82 | 93.3 | 97.1 | |
| MonodepthPost-processing=true, Data=S, Resolution=256 × 5122019.09 | 0.138 | 1.186 | 5.65 | 0.234 | 81.3 | 93 | 96.9 | |
| CCTrain=M, Resolution=832 × 2562021.12 | 0.14 | 1.07 | 5.326 | 0.217 | 82.6 | 94.1 | 97.5 | |
| Struct2depthTrain=M, Resolution=416 × 1282021.12 | 0.141 | 1.026 | 5.291 | 0.215 | 81.6 | 94.5 | 97.9 | |
| GeoNet (ResNet)Supervised=No, Dataset=K, Backbone=ResNet, Evaluation Cap=50m2018.03 | 0.147 | 0.936 | 4.348 | 0.218 | 81 | 94.1 | 97.7 | |
| Godard et al.Supervised=Pose, Dataset=K2018.03 | 0.148 | 1.344 | 5.927 | 0.247 | 80.3 | 92.2 | 96.4 | |
| Godard et al.Supervised=Pose2018.12 | 0.148 | 1.344 | 5.927 | 0.247 | 80.3 | 92.2 | 96.4 | |
| GeoNetTrain=M, Resolution=416 × 1282021.12 | 0.149 | 1.06 | 5.567 | 0.226 | 79.6 | 93.5 | 97.5 | |
| DF-NetTrain=M, Resolution=576 × 1602021.12 | 0.15 | 1.124 | 5.507 | 0.223 | 80.6 | 93.3 | 97.3 | |
| DDVOTrain=M, Resolution=416 × 1282021.12 | 0.151 | 1.257 | 5.583 | 0.228 | 81 | 93.6 | 97.4 | |
| GeoNet (ResNet)Supervised=No, Dataset=CS + K, Backbone=ResNet2018.03 | 0.153 | 1.328 | 5.737 | 0.232 | 80.2 | 93.4 | 97.2 | |
| GeoNet (ResNet)Supervised=No, Dataset=K, Backbone=ResNet2018.03 | 0.155 | 1.296 | 5.857 | 0.233 | 79.3 | 93.1 | 97.3 | |
| Yin et al.Supervised=No2018.12 | 0.155 | 1.296 | 5.857 | 0.233 | 79.3 | 93.1 | 97.3 | |
| GeoNet (VGG)Supervised=No, Dataset=K, Backbone=VGG, Evaluation Cap=50m2018.03 | 0.157 | 0.99 | 4.6 | 0.231 | 78.1 | 93.1 | 97.4 | |
| GeoNet (VGG)Supervised=No, Dataset=K, Backbone=VGG2018.03 | 0.164 | 1.303 | 6.09 | 0.247 | 76.5 | 91.9 | 96.8 | |
| Our T2Net, Dimage onlyDataset=vK(I+D) + K(I), cap=1-50m2018.08 | 0.168 | 1.199 | 4.674 | 0.243 | 77.2 | 91.2 | 96.6 | |
| Garg et al.Supervised=Pose, Dataset=K, Evaluation Cap=50m2018.03 | 0.169 | 1.08 | 5.104 | 0.273 | 74 | 90.4 | 96.2 | |
| Garg et al. [7] L12 Aug.8xDataset=K(L+R), cap=1-50m2018.08 | 0.169 | 1.08 | 5.104 | 0.273 | 74 | 90.4 | 96.2 | |
| Our full T2NetDataset=vK(I+D) + K(I), cap=1-50m2018.08 | 0.169 | 1.23 | 4.717 | 0.245 | 76.9 | 91.2 | 96.5 | |
| Zhou et al. updatedSupervised=No, Dataset=K2018.03 | 0.183 | 1.595 | 6.709 | 0.27 | 73.4 | 90.2 | 95.9 | |
| Zhou et al. updatedSupervised=No2018.12 | 0.183 | 1.595 | 6.709 | 0.27 | 73.4 | 90.2 | 95.9 | |
| SfMLeanerTrain=M, Resolution=416 × 1282021.12 | 0.183 | 1.595 | 6.709 | 0.27 | 73.4 | 90.2 | 95.9 | |
| Eigen et al. [4] FineDataset=K(I+D), cap=0-80m2018.08 | 0.19 | 1.515 | 7.156 | 0.27 | 69.2 | 89.9 | 96.7 | |
| Zhou et al.Supervised=No, Dataset=CS + K2018.03 | 0.198 | 1.836 | 6.565 | 0.275 | 71.8 | 90.1 | 96 | |
| Liu et al.Supervised=Depth, Dataset=K2018.03 | 0.202 | 1.614 | 6.523 | 0.275 | 67.8 | 89.5 | 96.5 | |
| Liu et al.Supervised=Depth2018.12 | 0.202 | 1.614 | 6.523 | 0.275 | 67.8 | 89.5 | 96.5 | |
| Eigen et al. FineSupervised=Depth, Dataset=K2018.03 | 0.203 | 1.548 | 6.307 | 0.282 | 70.2 | 89 | 95.8 | |
| Eigen et al. FineSupervised=Depth2018.12 | 0.203 | 1.548 | 6.307 | 0.282 | 70.2 | 89 | 95.7 | |
| Zhou et al.Supervised=No, Dataset=K2018.03 | 0.208 | 1.768 | 6.856 | 0.283 | 67.8 | 88.5 | 95.7 | |
| Eigen et al. CoarseSupervised=Depth, Dataset=K2018.03 | 0.214 | 1.605 | 6.563 | 0.292 | 67.3 | 88.4 | 95.7 | |
| Eigen et al. CoarseSupervised=Depth2018.12 | 0.214 | 1.605 | 6.653 | 0.292 | 67.3 | 88.4 | 95.7 | |
| Our T2Net, Dfeat onlyDataset=vK(I+D) + K(I), cap=1-50m2018.08 | 0.233 | 2.902 | 6.285 | 0.3 | 74.3 | 88 | 93.8 | |
| Our fr, all-syntheticDataset=vK(I+D), cap=1-50m2018.08 | 0.278 | 3.216 | 6.268 | 0.322 | 68.1 | 85.4 | 92.9 | |
| Baseline, train set meanDataset=vK(I+D), cap=1-50m2018.08 | 0.521 | 11.024 | 10.598 | 0.473 | 63.8 | 75.5 | 83.5 |