Video Depth Estimation on Sintel
76.3Delta Threshold Accuracy (1.25)PAGE-4D
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| PAGE-4DParams=1.26B, Align=scale&shift, Video Depth2025.10 | 76.3 | — | — | 21.2 | — | — | — | — | — | |
| PAGE-4DParams=1.26B, Align=Monocular Depth2025.10 | 74.2 | — | — | 24.2 | — | — | — | — | — | |
| pi^3Params=959M, Inference Speed Complexity=O(N^2), Alignment Protocol=joint scale-and-shift2026.03 | 73.5 | — | — | 20.6 | — | — | — | — | — | |
| TrackingWorld (DELTA)Depth Prior=Unidepth2025.12 | 73.3 | 0.218 | — | — | — | — | — | — | — | |
| Pi3alignment=per-sequence scale & shift2026.02 | 73.2 | 0.229 | — | — | — | — | — | — | — | |
| TrackingWorld (CoTrackerV3)Depth Prior=Unidepth2025.12 | 73.1 | 0.219 | — | — | — | — | — | — | — | |
| ZipMapParams=1.40B, Inference Speed Complexity=O(N), Alignment Protocol=joint scale-and-shift2026.03 | 73.1 | — | — | 19.8 | — | — | — | — | — | |
| VGGTalignment=per-sequence scale & shift2026.02 | 72.7 | 0.202 | — | — | — | — | — | — | — | |
| π³Align=scale & shift2026.02 | 72.6 | 0.21 | — | — | — | — | — | — | — | |
| pi^3Category=Joint D&P2026.01 | 72.6 | 0.21 | — | — | — | — | — | — | — | |
| Ours (DELTA)Category=Joint video depth & pose2025.12 | 72.6 | 0.222 | — | — | — | — | — | — | — | |
| TrajVGAlign=scale & shift2026.02 | 72.3 | 0.188 | — | — | — | — | — | — | — | |
| TrajVGAlign=scale2026.02 | 71.7 | 0.22 | — | — | — | — | — | — | — | |
| TurboVGGT2026.05 | 71.6 | — | — | 21.2 | — | — | — | — | — | |
| Ours (CoTrackerV3)Category=Joint video depth & pose2025.12 | 71.4 | 0.232 | — | — | — | — | — | — | — | |
| π^3Scale=Scale-Invariant, Online=false2026.05 | 71.4 | — | — | 0.217 | — | — | — | 3.911 | 1.44 | |
| DepthAnything3Scale=Scale-Invariant, Online=false2026.05 | 70 | — | — | 0.216 | — | — | — | 4.636 | 1.89 | |
| PAGE-4DParams=1.26B, Align=scale, Video Depth2025.10 | 69.9 | — | — | 35.7 | — | — | — | — | — | |
| DepthCrafteralignment=per-sequence scale & shift2026.02 | 69.7 | 0.292 | — | — | — | — | — | — | — | |
| DepthCrafterCategory=Video depth2026.01 | 69.7 | 0.292 | — | — | — | — | — | — | — | |
| DepthCrafterCategory=Depth Estimation model, Optim.=false, FF=true, Global Alignment (GA)=false, FPS=0.972026.01 | 69.7 | 0.292 | — | — | — | — | — | — | — | |
| DepthCrafterCategory=Video depth2025.12 | 69.7 | 0.292 | — | — | — | — | — | — | — | |
| ZipMap2026.05 | 69.5 | — | — | 24.8 | — | — | — | — | — | |
| V-DPMCategory=Joint D&P2026.01 | 69.4 | 0.247 | — | — | — | — | — | — | — | |
| VGGT2026.02 | 68.8 | 0.297 | — | — | — | — | — | — | — | |
| VGG-TAlignment=Per-sequence scale, Type=FA2025.12 | 68.8 | 0.297 | — | — | — | — | — | — | — | |
| VGGTType=FA2026.06 | 68.8 | — | — | 29.7 | — | — | — | — | — | |
| CogniMap3DCategory=Vision Foundation Model, Optim.=false, FF=true, Global Alignment (GA)=false, FPS=14.322026.01 | 68.6 | 0.295 | — | — | — | — | — | — | — | |
| VGGTParams=1.26B, Inference Speed Complexity=O(N^2), Alignment Protocol=joint scale-and-shift2026.03 | 68.3 | — | — | 22.6 | — | — | — | — | — | |
| VGGT2026.03 | 68.1 | — | — | 29.8 | — | — | — | — | — | |
| VGGTType=Dense-view2025.07 | 68.1 | — | — | 29.8 | — | — | — | — | — | |
| VGGTAlign=scale & shift2026.02 | 67.8 | 0.23 | — | — | — | — | — | — | — | |
| Pi32026.02 | 67.7 | 0.246 | — | — | — | — | — | — | — | |
| MoE3DType=FA2026.06 | 67.7 | — | — | 27.1 | — | — | — | — | — | |
| VGGT + Ours (GMM)Type=FA2026.06 | 67.4 | — | — | 24.1 | — | — | — | — | — | |
| 4RCalignment=per-sequence scale & shift2026.02 | 67 | 0.249 | — | — | — | — | — | — | — | |
| DA3 + Ours (GMM)Type=FA2026.06 | 67 | — | — | 22.3 | — | — | — | — | — | |
| DA3Type=FA2026.06 | 66.7 | — | — | 30.7 | — | — | — | — | — | |
| π³Align=scale2026.02 | 66.4 | 0.233 | — | — | — | — | — | — | — | |
| π³Processing Type=Offline2025.12 | 66.4 | 0.233 | — | — | — | — | — | — | — | |
| pi^3Params=959M2025.07 | 66.4 | — | — | 23.3 | — | — | — | — | — | |
| VGGTAlignment=Per-sequence scale, Online=false, Venue=CVPR’252026.03 | 66.1 | — | — | 0.287 | — | — | — | — | — | |
| StreamVGGTProcessing Type=Online2025.12 | 65.7 | 0.323 | — | — | — | — | — | — | — | |
| StreamVGGTAlignment=Per-sequence scale, Online=true, Venue=ArXiv’252026.03 | 65.7 | — | — | 0.323 | — | — | — | — | — | |
| StreamVGGTType=Streaming2025.07 | 65.7 | — | — | 32.3 | — | — | — | — | — | |
| VGGTScale=Scale-Invariant, Online=false2026.05 | 65.7 | — | — | 0.231 | — | — | — | 5.517 | 3.56 | |
| UNITScale=Scale-Invariant, G=N, Online=false2026.05 | 65.4 | — | — | 0.215 | — | — | — | 4.78 | 3.11 | |
| VGGT2026.05 | 64.6 | — | — | 30 | — | — | — | — | — | |
| IVGTdecoding=direct2026.05 | 64.6 | — | — | 29.5 | — | — | — | — | — | |
| MoRetype=FA2026.03 | 64.5 | — | — | 0.335 | — | — | — | — | — | |
| Align3R (Depth Pro)Category=Joint video depth & pose2025.12 | 64.1 | 0.263 | — | — | — | — | — | — | — | |
| VGGTParams=1.26B, Align=scale&shift, Video Depth2025.10 | 63.9 | — | — | 26.1 | — | — | — | — | — | |
| Complet4R2026.03 | 63.9 | — | — | 35.3 | — | — | — | — | — | |
| SparseVGGT2026.05 | 63.9 | — | — | 30.4 | — | — | — | — | — | |
| VGGTAlign=scale2026.02 | 63.8 | 0.299 | — | — | — | — | — | — | — | |
| VGGTProcessing Type=Offline2025.12 | 63.8 | 0.299 | — | — | — | — | — | — | — | |
| ZipMapStreaming=true, Alignment=scale-only2026.03 | 63.8 | — | — | 27.3 | — | — | — | — | — | |
| VGGTParams=1.26B2025.07 | 63.8 | — | — | 29.9 | — | — | — | — | — | |
| VGGT2026.05 | 63.8 | — | — | 29.9 | — | — | — | — | — | |
| MoRetype=Streaming2026.03 | 63.7 | — | — | 0.254 | — | — | — | — | — | |
| Stream3Rtype=Streaming2026.03 | 63.2 | — | — | 0.397 | — | — | — | — | — | |
| DELTADepth Prior=Unidepth2025.12 | 63.1 | 0.636 | — | — | — | — | — | — | — | |
| UnidepthCategory=Single-frame depth2025.12 | 63 | 0.473 | — | — | — | — | — | — | — | |
| FastVGGT2026.05 | 63 | — | — | 30.7 | — | — | — | — | — | |
| VGGTParams=1.26B, Align=Monocular Depth2025.10 | 62.9 | — | — | 29.2 | — | — | — | — | — | |
| Pi3MOS-SLAMProcessing Type=Online2025.12 | 62.5 | 0.287 | — | — | — | — | — | — | — | |
| VGGTCategory=Vision Foundation Model, Optim.=false, FF=true, Global Alignment (GA)=false, FPS=21.52026.01 | 62.4 | 0.299 | — | — | — | — | — | — | — | |
| 4RC2026.02 | 62.2 | 0.311 | — | — | — | — | — | — | — | |
| UNITScale=Scale-Invariant, G=1, Online=true2026.05 | 60.9 | — | — | 0.253 | — | — | — | 5.13 | 1.22 | |
| AetherAlign=scale & shift2026.02 | 60.4 | 0.314 | — | — | — | — | — | — | — | |
| StreamVGGTScale=Scale-Invariant, Online=true2026.05 | 59.7 | — | — | 0.27 | — | — | — | 5.511 | 2.11 | |
| MASt3R-GAalignment=per-sequence scale & shift2026.02 | 59.4 | 0.327 | — | — | — | — | — | — | — | |
| MASt3R-GACategory=Vision Foundation Model, Optim.=true, FF=false, Global Alignment (GA)=true, FPS=0.312026.01 | 59.4 | 0.327 | — | — | — | — | — | — | — | |
| Depth Anything V2Category=Single-frame depth2025.12 | 59.2 | 0.348 | — | — | — | — | — | — | — | |
| StreamVGGTtype=Streaming2026.03 | 59.1 | — | — | 0.698 | — | — | — | — | — | |
| MonST3R-GAalignment=per-sequence scale & shift2026.02 | 59 | 0.333 | — | — | — | — | — | — | — | |
| MonST3R-GACategory=Vision Foundation Model, Optim.=true, FF=false, Global Alignment (GA)=true, FPS=0.352026.01 | 59 | 0.333 | — | — | — | — | — | — | — | |
| MonST3RCategory=Joint video depth & pose2025.12 | 58.6 | 0.335 | — | — | — | — | — | — | — | |
| MonST3RCategory=Joint D&P2026.01 | 58.5 | 0.335 | — | — | — | — | — | — | — | |
| VGGTtype=FA2026.03 | 58.4 | — | — | 0.387 | — | — | — | — | — | |
| OmniStreamParam.=400M2026.03 | 58.3 | — | — | 31.4 | — | — | — | — | — | |
| IVGTdecoding=from render2026.05 | 58.2 | — | — | 54.2 | — | — | — | — | — | |
| VGG-T32026.05 | 58.1 | — | — | 34.5 | — | — | — | — | — | |
| DPMCategory=Joint D&P2026.01 | 58 | 0.311 | — | — | — | — | — | — | — | |
| DA3 + Ours (LMM)Type=FA2026.06 | 57.9 | — | — | 33.3 | — | — | — | — | — | |
| MASt3RParams=689M, Align=Monocular Depth2025.10 | 56.9 | — | — | 41.3 | — | — | — | — | — | |
| TTT3RParams=793M, Inference Speed Complexity=O(N), Alignment Protocol=joint scale-and-shift2026.03 | 56.6 | — | — | 50.8 | — | — | — | — | — | |
| MapAnythingScale=Scale-Invariant, Online=false2026.05 | 56.5 | — | — | 0.463 | — | — | — | 5.73 | 5 | |
| Depth ProCategory=Single-frame depth2025.12 | 55.9 | 0.418 | — | — | — | — | — | — | — | |
| Easi3RAlignment=Per-sequence scale, Online=false, Venue=ICCV’252026.03 | 55.9 | — | — | 0.377 | — | — | — | — | — | |
| MonST3R2026.02 | 55.8 | 0.378 | — | — | — | — | — | — | — | |
| CUT3RAlign=scale & shift2026.02 | 55.8 | 0.534 | — | — | — | — | — | — | — | |
| MonST3R-GAAlignment=Per-sequence scale, Type=Optim2025.12 | 55.8 | 0.378 | — | — | — | — | — | — | — | |
| MonSTR3RAlignment=Per-sequence scale, Online=false, Venue=ICLR’252026.03 | 55.8 | — | — | 0.378 | — | — | — | — | — | |
| CUT3RParams=793M, Align=scale&shift, Video Depth2025.10 | 55.8 | — | — | 53.4 | — | — | — | — | — | |
| MonST3R-GAType=Pair-wise2025.07 | 55.8 | — | — | 37.8 | — | — | — | — | — | |
| CUT3Ralignment=per-sequence scale & shift2026.02 | 55.7 | 0.54 | — | — | — | — | — | — | — | |
| CUT3RCategory=Vision Foundation Model, Optim.=false, FF=true, Global Alignment (GA)=false, FPS=16.582026.01 | 55.7 | 0.454 | — | — | — | — | — | — | — | |
| Depth-Anything-V2alignment=per-sequence scale & shift2026.02 | 55.4 | 0.367 | — | — | — | — | — | — | — | |
| DepthAnythingV2Category=1-frame2026.01 | 55.4 | 0.367 | — | — | — | — | — | — | — |