Point Tracking on DAVIS TAP-Vid
69.4Average Jaccard (AJ)LocoTrack-B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LocoTrack-BInput Resolution=384x512, Query Mode=Strided, Throughput=4196.36 points/sec2024.07 | 69.4 | 81.3 | 88.6 | — | |
| LocoTrack-SInput Resolution=384x512, Query Mode=Strided, Throughput=6820.57 points/sec2024.07 | 68.4 | 80.4 | 87.5 | — | |
| Track-On-RTraining Strategy=Real-World Fine-Tuning2026.03 | 68.1 | — | 92.5 | 80.3 | |
| LocoTrack-BInput Resolution=256x256, Query Mode=Strided, Throughput=4358.96 points/sec2024.07 | 67.8 | 79.6 | 89.9 | — | |
| Track-On2Training Strategy=Synthetic Pretraining2026.03 | 67 | — | 92 | 79.9 | |
| LocoTrack-SInput Resolution=256x256, Query Mode=Strided, Throughput=7244.47 points/sec2024.07 | 66.9 | 78.8 | 88.9 | — | |
| FlowTrackInput Resolution=384x512, Query Mode=Strided2024.07 | 66 | 79.8 | 87.2 | — | |
| CoTrackerInput Resolution=384x512, Query Mode=Strided, Throughput=1146.79 points/sec2024.07 | 65.9 | 79.4 | 89.9 | — | |
| BootsTAPNext-BTraining Strategy=Real-World Fine-Tuning2026.03 | 65.2 | — | 91.2 | 78.5 | |
| LocoTrackTraining Dataset=Kub., Steps=400k, Sequence Length=24-48, Batch Size=82025.12 | 64.8 | 77.4 | 86.2 | — | |
| CoTracker3†Training Dataset=Kub. + 15k, Steps=65k, Sequence Length=60-80, Batch Size=322025.12 | 64.8 | 76.8 | 91.7 | — | |
| Anthro-LocoTrackTraining Strategy=Real-World Fine-Tuning2026.03 | 64.8 | — | 89.1 | 77.3 | |
| CoTracker3 (Window)Training Strategy=Synthetic Pretraining2026.03 | 64.5 | — | 89.7 | 76.7 | |
| TAPNextSupervision=Supervised2025.12 | 64.48 | — | 91.71 | 77.29 | |
| CoTracker3Supervision=Supervised2025.12 | 64.45 | — | 90.9 | 77.13 | |
| CoTracker3 (Window)Training Strategy=Real-World Fine-Tuning2026.03 | 63.8 | — | 90.2 | 76.3 | |
| AllTrackerTraining Strategy=Real-World Fine-Tuning, Leverages additional optical-flow=true2026.03 | 63.7 | — | 88.7 | 77 | |
| TAPTRv2Training Dataset=Kub., Steps=44k, Sequence Length=-, Batch Size=322025.12 | 63 | 76.1 | 91.1 | — | |
| LocoTrackTraining Strategy=Synthetic Pretraining2026.03 | 63 | — | 87.2 | 75.3 | |
| TAPIRTuning Dataset=MOVi-E, Version=Open-source tuned2023.06 | 62.9 | — | — | — | |
| DiTrackerTraining Dataset=Kub., Steps=36k, Sequence Length=46, Batch Size=42025.12 | 62.7 | 77.5 | 85.2 | — | |
| TAPIRTuning Dataset=Panning Kubric, Version=Open-source tuned2023.06 | 62.4 | — | — | — | |
| TAPNext-BTraining Strategy=Synthetic Pretraining2026.03 | 62.4 | — | 90.5 | 76.6 | |
| DINTR2024.10 | 62.3 | 74.6 | 88.9 | — | |
| BootsTAPIRTraining Dataset=Kub. + 15M, Steps=200k, Sequence Length=24, Batch Size=1,5362025.12 | 61.4 | 73.6 | 88.7 | — | |
| BootsTAPIRTraining Strategy=Real-World Fine-Tuning2026.03 | 61.4 | — | 88.7 | 73.6 | |
| TAPIRInput Resolution=256x256, Query Mode=Strided, Throughput=2097.32 points/sec2024.07 | 61.3 | 73.6 | 88.8 | — | |
| TAPIR2024.10 | 59.8 | 72.3 | 87.6 | — | |
| CoTracker3Training Dataset=Kub., Steps=50k, Sequence Length=60, Batch Size=322025.12 | 57.8 | 74.9 | 80.4 | — | |
| TAPIRSupervision=Supervised2025.12 | 57.01 | — | 86.33 | 69.34 | |
| TAPIRTraining Dataset=Kub., Steps=50k, Sequence Length=24, Batch Size=42025.12 | 56.2 | 70 | 86.5 | — | |
| TAPIRTraining Strategy=Synthetic Pretraining2026.03 | 56.2 | — | 86.5 | 70 | |
| HeFTSupervision=Zero-Shot, Backbone=Cosmos-Predict2-2B2025.12 | 48.61 | — | 82.47 | 63.5 | |
| Opt-CWMSupervision=Self-Supervised2025.12 | 47.53 | — | 80.87 | 64.83 | |
| Point-PromptingSupervision=Zero-Shot2025.12 | 42.21 | — | 82.9 | 57.29 | |
| PIPsInput Resolution=384x512, Query Mode=Strided, Throughput=46.43 points/sec2024.07 | 42 | 59.4 | 82.1 | — | |
| PIPs2024.10 | 42 | 59.4 | 82.1 | — | |
| HeFTSupervision=Zero-Shot, Backbone=Wan2.1-1.3B2025.12 | 41.2 | — | 80 | 54.9 | |
| TAP-NetInput Resolution=256x256, Query Mode=Strided, Throughput=29535.98 points/sec2024.07 | 38.4 | 53.1 | 82.3 | — | |
| TAP-Net2024.10 | 38.4 | 53.1 | 82.3 | — | |
| GMRWSupervision=Self-Supervised2025.12 | 36.47 | — | 76.36 | 54.659 | |
| COTR2024.10 | 35.4 | 51.3 | 80.2 | — | |
| Kubric-VFS-LikeInput Resolution=256x256, Query Mode=Strided2024.07 | 33.1 | 48.5 | 79.4 | — | |
| Kubric-VFS-Like2024.10 | 33.1 | 48.5 | 79.4 | — | |
| TAP-NetSupervision=Supervised2025.12 | 32.05 | — | 77.35 | 48.42 | |
| RAFTInput Resolution=256x256, Query Mode=Strided, Throughput=23405.71 points/sec2024.07 | 30 | 46.3 | 79.6 | — | |
| RAFT2024.10 | 30 | 46.3 | 79.6 | — | |
| SD-DINOSupervision=Zero-Shot2025.12 | 29.68 | — | 69.71 | 50.54 | |
| RAFTSupervision=Supervised2025.12 | 25.06 | — | 74.69 | 43.64 | |
| HeFTSupervision=Zero-Shot, Backbone=CogvideoX-2B2025.12 | 24.72 | — | 71.87 | 31.18 | |
| DIFTSupervision=Zero-Shot2025.12 | 21.51 | — | 69.71 | 39.55 | |
| DINOv2+NNSupervision=Zero-Shot2025.12 | 15.19 | — | 61.81 | 31.19 | |
| AccFlowzero-shot=true, input resolution=384 × 512, trained on=optical flow datasets exclusively2026.03 | — | 23.5 | — | — | |
| DiffTrackSupervision=Zero-Shot2025.12 | — | — | — | 46.21 | |
| MegaFlowzero-shot=true, input resolution=384 × 512, trained on=optical flow datasets exclusively2026.03 | — | 65.6 | — | — | |
| MemFlow-Tzero-shot=true, input resolution=384 × 512, trained on=optical flow datasets exclusively2026.03 | — | 61.7 | — | — | |
| RAFTzero-shot=true, input resolution=384 × 512, trained on=optical flow datasets exclusively2026.03 | — | 48.5 | — | — | |
| SEA-RAFTzero-shot=true, input resolution=384 × 512, trained on=optical flow datasets exclusively2026.03 | — | 48.7 | — | — | |
| WAFT-DINOv3-a2zero-shot=true, input resolution=384 × 512, trained on=optical flow datasets exclusively2026.03 | — | 53.9 | — | — |