Natural Language Video Localization on ActivityNet Caption (test)
43.6IoU @ 0.5ExCL
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ExCLvariation=*, note=missing some videos2020.04 | 43.6 | 63 | 24.1 | — | |
| VSLNet2020.04 | 43.22 | 63.16 | 26.16 | 43.19 | |
| DEBUG2020.04 | 39.72 | 55.91 | — | 39.51 | |
| VSLBase2020.04 | 39.52 | 58.18 | 23.21 | 40.56 | |
| RWM-RL2020.04 | 36.9 | — | — | — | |
| ABLR2020.04 | 36.79 | 55.67 | — | 36.99 | |
| LT-ZVGArchitecture Type=Specialist model, # Train Samples=0, Setup=Zero-shot, Trained on raw videos (no annotations)=true2026.01 | 32.6 | 47.6 | 15.4 | 31.8 | |
| HawkEyeArchitecture Type=MLLM-based model, # Train Samples=715K, Setup=Zero-shot2026.01 | 29.3 | 49.1 | 10.7 | 32.7 | |
| TGN2020.04 | 28.47 | 45.51 | — | — | |
| VideoChat2Architecture Type=MLLM-based model, # Train Samples=2M, Setup=Zero-shot2026.01 | 27.8 | 40.8 | 9.3 | 27.9 | |
| TimeLLMArchitecture Type=MLLM-based model, # Train Samples=170K, Setup=Zero-shot2026.01 | 27.8 | 44 | 14.3 | 30.4 | |
| QSPN2020.04 | 27.7 | 45.3 | 13.6 | — | |
| VIRTUE-Embed 7BArchitecture Type=MLLM-based model, # Train Samples=0, Setup=Zero-shot2026.01 | 26.7 | 47.2 | 14.5 | 33.4 | |
| MomenterArchitecture Type=MLLM-based model, # Train Samples=10M, Setup=Zero-shot2026.01 | 23 | 42.9 | 12.4 | 29.3 | |
| ChatVTGArchitecture Type=MLLM-based model, # Train Samples=100K, Setup=Zero-shot2026.01 | 22.5 | 40.7 | 9.4 | 27.2 | |
| SeViLAArchitecture Type=Specialist model, # Train Samples=129M, Setup=Zero-shot2026.01 | 19 | 31.6 | 10.1 | 23 |