Temporal Video Grounding on ActivityNet Captions (test)
84.73Recall@IoU>0.5UniversalVTG
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| UniversalVTGShared Weights=true2026.04 | 84.73 | — | — | 63.64 | 74.19 | — | |
| Choice 3 rounds UpboundZero-shot=true2024.03 | 84.6 | 71.9 | 91.5 | 68.4 | — | — | |
| SnAGShared Weights=false2026.04 | 81.71 | — | — | 63.41 | 72.56 | — | |
| UnLoc-LShared Weights=false2026.04 | 79.2 | — | — | 61.3 | 70.25 | — | |
| 2D-TANShared Weights=false2026.04 | 78.8 | — | — | 60.85 | 69.83 | — | |
| Mr.BLIPModel Category=Task-Specific Models2025.10 | 53.9 | — | — | 35.5 | — | — | |
| BM-DETRModel Category=Task-Specific Models2025.10 | 49.6 | — | — | 30.6 | — | — | |
| SOTA SpecialistFine-tuned=true2024.03 | 43.8 | 44.2 | 63.2 | 27.1 | — | — | |
| 2D-TANYear=2020, Setup=FS2024.03 | 43.41 | 42.45 | 60.32 | 25.04 | — | — | |
| TimeChatModel Category=Video-LLMs, Evaluation Protocol=Temporally Conditioned Attention Sharpening (TCAS)2025.10 | 41.2 | — | — | 24.9 | — | — | |
| TimeChatModel Category=Video-LLMs, Evaluation Protocol=Supervised Fine-Tuning (SFT)2025.10 | 41 | — | — | 23.7 | — | — | |
| Huang et al.Year=2023, Setup=WS2024.03 | 36.91 | 41.02 | 58.07 | — | — | — | |
| DRIFT (ET-Chat)# Train=100K, Protocol=zero-shot2026.06 | 35.8 | 37.3 | 57.9 | 14.1 | — | — | |
| VideoChat2Fine-tuned=true, Implementation=our impl.2024.03 | 34.7 | 38.9 | 55.5 | 17.7 | — | — | |
| HawkEyeFine-tuned=true2024.03 | 34.7 | 39.1 | 55.9 | 17.9 | — | — | |
| HawkEyeModel Category=Video-LLMs2025.10 | 34.7 | — | — | 17.7 | — | — | |
| CNMYear=2022, Setup=WS2024.03 | 33.33 | 37.55 | 55.68 | 13.29 | — | — | |
| ED-VTG# Train=136K, Protocol=zero-shot2026.06 | 33.1 | 35.2 | 52.1 | 16 | — | — | |
| LT-ZVGProtocol=zero-shot2026.06 | 32.6 | 31.8 | 47.6 | 15.4 | — | — | |
| Kim et al.Year=2023, Setup=US2024.03 | 32.59 | 31.85 | 47.61 | 15.42 | — | — | |
| CPLYear=2022, Setup=WS2024.03 | 31.37 | 36.65 | 55.73 | 10.68 | — | — | |
| PZVMRYear=2022, Setup=US2024.03 | 31.26 | 30.35 | 45.73 | 17.84 | — | — | |
| PSVLProtocol=zero-shot2026.06 | 30.1 | 29.6 | 44.7 | 14.7 | — | — | |
| PSVLYear=2021, Setup=US2024.03 | 30.08 | 29.62 | 44.74 | 14.74 | — | — | |
| HawkEyeZero-shot=true2024.03 | 29.3 | 32.7 | 49.1 | 10.7 | — | — | |
| HawkEye# Train=715K, Protocol=zero-shot2026.06 | 29.3 | 32.7 | 49.1 | 10.7 | — | — | |
| VDIYear=2023, Setup=FS2024.03 | 28.76 | — | 48.09 | — | — | — | |
| VideoChat2Zero-shot=true, implementation=our implementation2024.03 | 28.7 | 28.2 | 41.7 | 9.4 | — | — | |
| VTG-GPTYear=2023, Setup=ZS2024.03 | 28.25 | 30.49 | 47.13 | 12.84 | — | — | |
| DSCNetYear=2022, Setup=US2024.03 | 28.16 | — | 47.29 | — | — | — | |
| TimeChatModel Category=Video-LLMs2025.10 | 28 | — | — | 15.8 | — | — | |
| Luo et al.Year=2023, Setup=ZS2024.03 | 27.9 | 32.37 | 48.28 | 11.57 | — | — | |
| VideoChat2Zero-shot=true2024.03 | 27.8 | 27.9 | 40.8 | 9.3 | — | — | |
| VTimeLLM-7BZero-shot=true2024.03 | 27.8 | 30.4 | 44 | 14.3 | — | — | |
| VideoChat2# Train=2M, Protocol=zero-shot2026.06 | 27.8 | 27.9 | 40.8 | 9.3 | — | — | |
| VTimeLLM# Train=170K, Protocol=zero-shot2026.06 | 27.8 | 30.4 | 44 | 14.3 | — | — | |
| Gao et al.Year=2021, Setup=US2024.03 | 26.38 | — | 46.15 | 11.64 | — | — | |
| Momentor# Train=10M, Protocol=zero-shot2026.06 | 23 | 29.3 | 42.9 | 12.4 | — | — | |
| ChatVTG# Train=100K, Protocol=zero-shot2026.06 | 22.5 | 27.2 | 40.7 | 9.4 | — | — | |
| SeViLA LocalizerZero-shot=true2024.03 | 19 | 23 | 31.6 | 10.1 | — | — | |
| SeViLA# Train=129M, Protocol=zero-shot2026.06 | 19 | 23 | 31.6 | 10.1 | — | — | |
| RandomZero-shot=true2024.03 | 15.1 | 23 | 29 | 6.1 | — | — | |
| Valley# Train=100K, Protocol=zero-shot2026.06 | 13.7 | 21.9 | 30.6 | 8.1 | — | — | |
| ET-Chat# Train=164K, Protocol=zero-shot2026.06 | 12.7 | 18.9 | 24.1 | 6.2 | — | — | |
| Video-LLaMA# Train=2.7M, Protocol=zero-shot2026.06 | 10.8 | 16.5 | 21.9 | 4.9 | — | — | |
| Video-ChatGPT# Train=100K, Protocol=zero-shot2026.06 | 10.6 | 14.2 | 19.5 | 4.8 | — | — | |
| ActivityNet-Captions ExpertTraining Dataset=ActivityNet-Captions2026.04 | — | — | — | — | — | 39.52 | |
| Charades-STA ExpertTraining Dataset=Charades-STA2026.04 | — | — | — | — | — | 9.35 | |
| NLQ ExpertTraining Dataset=Ego4D-NLQ2026.04 | — | — | — | — | — | 13.16 | |
| Prior Unified SOTATraining Dataset=Unified2026.04 | — | — | — | — | — | 31.98 | |
| TACoS ExpertTraining Dataset=TACoS2026.04 | — | — | — | — | — | 10.78 | |
| UniversalVTGTraining Dataset=Unified2026.04 | — | — | — | — | — | 39.39 |