Video Depth Estimation on NYUDV2 (test)
95delta1NVDS+
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| NVDS+Type=Learning Based, Backbone=DPT-Large2023.07 | 95 | 0.072 | 33.9 | — | — | — | — | |
| NVDS+Type=Learning Based, Backbone=MiDaS-v2.1-Large2023.07 | 94.1 | 0.076 | 34.7 | — | — | — | — | |
| DPT-LargeType=Single Image2023.07 | 92.8 | 0.084 | 81.1 | — | — | — | — | |
| DeepV2DType=Learning Based2023.07 | 92.4 | 0.082 | 40.2 | — | — | — | — | |
| VITAType=Learning Based2023.07 | 92.2 | 0.092 | 38.5 | — | — | — | — | |
| MAMOType=Learning Based2023.07 | 91.9 | 0.094 | — | — | — | — | — | |
| MiDaS-v2.1-LargeType=Single Image2023.07 | 91 | 0.095 | 86.2 | — | — | — | — | |
| Robust-CVDType=Test-time Training2023.07 | 88.6 | 0.103 | 39.4 | — | — | — | — | |
| Cao et al.Type=Learning Based2023.07 | 83.5 | 0.131 | — | — | — | — | — | |
| ST-CLSTMType=Learning Based2023.07 | 83.3 | 0.131 | 64.5 | — | — | — | — | |
| FMNetType=Learning Based2023.07 | 83.2 | 0.134 | 38.7 | — | — | — | — | |
| WSVDType=Learning Based2023.07 | 76.8 | 0.164 | 68.3 | — | — | — | — | |
| DAv2-BType=Image to Depth2026.04 | — | — | — | 80 | 12.5 | 12.5 | — | |
| DAv2-B + NVDSType=Post-processing2026.04 | — | — | — | 3 | 367.8 | 12.5 | 355.3 | |
| DAv2-B + VDPPType=Post-processing2026.04 | — | — | — | 47 | 21.1 | 12.5 | 8.6 | |
| DAv2-LType=Image to Depth2026.04 | — | — | — | 48 | 20.8 | 20.8 | — | |
| DAv2-L + NVDSType=Post-processing2026.04 | — | — | — | 3 | 376.1 | 20.8 | 355.3 | |
| DAv2-L + VDPPType=Post-processing2026.04 | — | — | — | 34 | 29.4 | 20.8 | 8.6 | |
| DPT-LType=Image to Depth2026.04 | — | — | — | 82 | 12.2 | 12.2 | — | |
| DPT-L + NVDSType=Post-processing2026.04 | — | — | — | 3 | 367.5 | 12.2 | 355.3 | |
| DPT-L + VDPPType=Post-processing2026.04 | — | — | — | 48 | 20.8 | 12.2 | 8.6 | |
| VDA-LType=End-to-end2026.04 | — | — | — | 1 | 669.1 | — | — | |
| VDA-SType=End-to-end2026.04 | — | — | — | 14 | 72.7 | — | — |