Metric Depth Estimation on iBIMS-1
97.8Delta 1 Threshold AccUDv2 + BokehDepth
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| UDv2 + BokehDepthProtocol=Zero-shot, Backbone=UDv22025.12 | 97.8 | 0.039 | — | — | |
| DAv2 + BokehDepthProtocol=Zero-shot, Backbone=DAv22025.12 | 96 | 0.064 | — | — | |
| Ours-4bParameters=4b2026.05 | 96 | — | — | — | |
| VLM3-4bModel Type=VLM Trained on Metric Depth Estimation2026.05 | 96 | — | — | — | |
| UDv2Protocol=Zero-shot2025.12 | 94.5 | 0.082 | — | — | |
| UnidepthV22026.05 | 94.5 | — | — | — | |
| Metric3Dv2Protocol=Zero-shot2025.12 | 94.1 | 0.1 | — | — | |
| DAv2Protocol=Zero-shot2025.12 | 93.8 | 0.08 | — | — | |
| MoGe-22026.05 | 92.4 | — | — | — | |
| DepthLM-7bModel Type=VLM Trained on Metric Depth Estimation2026.05 | 92 | — | — | — | |
| DepthLM-3bModel Type=VLM Trained on Metric Depth Estimation2026.05 | 89 | — | — | — | |
| DepthProProtocol=Zero-shot2025.12 | 82.3 | 0.159 | — | — | |
| Depth Pro2026.05 | 82.3 | — | — | — | |
| Metric3D2026.05 | 79.7 | — | — | — | |
| TR2MBackbone=ViT-Small, Train Images=102K, Zero-shot=true2025.06 | 73.6 | 0.154 | — | 2.4 | |
| UniK3D-V-LBackbone=ViT-Large, Train Images=8M, Zero-shot=true2025.06 | 73.5 | 0.186 | — | 3.75 | |
| Depth AnythingZero-shot=true, Training set=NYUv22024.01 | 71.4 | 0.15 | — | — | |
| DA SingleBackbone=ViT-Large, Train Images=-, Zero-shot=true2025.06 | 71.4 | 0.15 | — | 2.8 | |
| DepthAnything2026.05 | 71.4 | — | — | — | |
| DA MixBackbone=ViT-Large, Train Images=102K, Zero-shot=true2025.06 | 70.5 | 0.158 | — | 5 | |
| Metric3Dv22026.05 | 68.4 | — | — | — | |
| ZoeDepthProtocol=Zero-shot2025.12 | 67.2 | 0.174 | — | — | |
| ZoeDepthZero-shot=true, Training set=NYUv22024.01 | 65.6 | 0.169 | — | — | |
| ZoeDepthBackbone=BeiT384-Large, Train Images=-, Zero-shot=true2025.06 | 65.6 | 0.169 | — | 4.7 | |
| Seed1.5-VLModel Type=VLM Trained on Metric Depth Estimation2026.05 | 62.7 | — | — | — | |
| ZoeDepth2026.05 | 58 | — | — | — | |
| Gemini-2.5-ProModel Type=VLM2026.05 | 46.6 | — | — | — | |
| RSABackbone=ViT-Small, Train Images=102K, Zero-shot=true2025.06 | 45 | 0.266 | — | 5.8 | |
| UniK3D-V-SBackbone=ViT-Small, Train Images=8M, Zero-shot=true2025.06 | 41.5 | 0.504 | — | 6 | |
| GPT-5Model Type=VLM2026.05 | 30.7 | — | — | — | |
| SpatialRGPT-8BModel Type=Spatial VLM2026.05 | 24 | — | — | — | |
| Unidepth-V-LBackbone=ViT-Large, Train Images=3M, Zero-shot=true2025.06 | 21.7 | 0.384 | — | 4 | |
| Qwen2.5-VL-72bModel Type=VLM2026.05 | 21.2 | — | — | — | |
| SpaceLLaVA-13BModel Type=Spatial VLM2026.05 | 20.8 | — | — | — | |
| Unidepth2026.05 | 15.7 | — | — | — | |
| Qwen3-VL-32bModel Type=VLM2026.05 | 12.2 | — | — | — | |
| Qwen3-VL-4bModel Type=VLM2026.05 | 8 | — | — | — | |
| AdabinsPre-training Dataset=NYUv2, Zero-shot Protocol=True2023.07 | — | 0.212 | 0.901 | — | |
| NewCRFsPre-training Dataset=NYUv2, Zero-shot Protocol=True2023.07 | — | 0.206 | 0.861 | — | |
| Ours_CSTM_imageZero-shot Protocol=True2023.07 | — | 0.144 | 0.646 | — | |
| Ours_CSTM_labelZero-shot Protocol=True2023.07 | — | 0.16 | 0.521 | — |