Temporal Video Grounding on Charades-STA (test)
97Recall@IoU=0.5Choice 3 rounds Upbound
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Choice 3 rounds UpboundZero-shot=true2024.03 | 97 | 74.8 | 100 | 69.2 | |
| Invert4TVGType=RL, Size=7B, FT=true, extra_pretraining=false2025.08 | 72.5 | — | 83 | 51.4 | |
| Time-R1Type=RL, Size=7B, FT=true, extra_pretraining=true2025.08 | 72.2 | — | 82.8 | 50.1 | |
| InternVideo2-S2-6BFeature=InternVideo2-S2-6B, Evaluation Protocol=Finetuned, Grounding head=CG-DETR2024.03 | 70.03 | 58.79 | 79.7 | 48.95 | |
| Mr.BLIPModel Category=Task-Specific Models2025.10 | 69.3 | — | — | 49.2 | |
| Invert4TVGType=RL, Size=3B, FT=true, extra_pretraining=false2025.08 | 69 | — | 80.8 | 44 | |
| EaTRType=VLP, Size=N/A, FT=true, extra_pretraining=true2025.08 | 68.4 | — | — | 44.9 | |
| InternVideo2-S2-1BFeature=InternVideo2-S2-1B, Evaluation Protocol=Finetuned, Grounding head=CG-DETR2024.03 | 68.36 | 57.12 | 78.41 | 45.03 | |
| VideoChat-TLLM Size=7B, Training Mode=Fine-tuned2024.10 | 67.1 | — | 79.4 | 43 | |
| TimeSuiteType=SFT, Size=7B, FT=true, extra_pretraining=true2025.08 | 67.1 | — | 79.4 | 43 | |
| SnAGType=VLP, Size=N/A, FT=true, extra_pretraining=true2025.08 | 64.6 | — | — | 46.2 | |
| Time-R1Type=RL, Size=3B, FT=true, extra_pretraining=true2025.08 | 64.1 | — | 78.7 | 36.9 | |
| UnLoc-LTraining Mode=Fine-tuned, Model Category=Expert model2024.10 | 60.8 | — | — | 38.4 | |
| Time-R1Type=RL, Size=7B, FT=false, extra_pretraining=false2025.08 | 60.8 | — | 78.1 | 35.3 | |
| TimeChatModel Category=Video-LLMs, Evaluation Protocol=Temporally Conditioned Attention Sharpening (TCAS)2025.10 | 60.2 | — | — | 37.6 | |
| BAM-DETR2023.11 | 59.95 | 52.33 | 72.93 | 39.38 | |
| BM-DETRModel Category=Task-Specific Models2025.10 | 59.4 | — | — | 38.3 | |
| CLIP+SlowFastFeature=CLIP+SlowFast, Evaluation Protocol=Finetuned, Grounding head=CG-DETR2024.03 | 58.44 | 50.13 | 70.43 | 36.34 | |
| TimeChatModel Category=Video-LLMs, Evaluation Protocol=Supervised Fine-Tuning (SFT)2025.10 | 58.4 | — | — | 34.7 | |
| HawkEyeFine-tuned=true2024.03 | 58.3 | 49.3 | 72.5 | 28.8 | |
| HawkEyeLLM Size=7B, Training Mode=Fine-tuned2024.10 | 58.3 | — | 72.5 | 28.8 | |
| HawkEyeType=SFT, Size=7B, FT=true, extra_pretraining=true2025.08 | 58.3 | — | 72.5 | 28.8 | |
| HawkEyeModel Category=Video-LLMs2025.10 | 58.3 | — | — | 28.8 | |
| UniVTG2023.11 | 58.01 | 50.1 | 70.81 | 35.65 | |
| QD-DETR2023.11 | 57.31 | — | — | 32.55 | |
| SOTA SpecialistFine-tuned=true2024.03 | 57.3 | — | — | 32.5 | |
| QD-DETRTraining Mode=Fine-tuned, Model Category=Expert model2024.10 | 57.3 | — | — | 32.6 | |
| VTG-LLMModel Category=Video-LLMs2025.10 | 57.2 | — | — | 33.4 | |
| VideoChat2Fine-tuned=true, Implementation=our impl.2024.03 | 56.4 | 48.3 | 71.8 | 27.7 | |
| MomentDiff2023.11 | 55.57 | — | — | 32.42 | |
| VideoChat-FlashType=SFT, Size=7B, FT=false, extra_pretraining=false2025.08 | 53.1 | — | 74.5 | 27.6 | |
| Time-R1Type=RL, Size=3B, FT=false, extra_pretraining=false2025.08 | 53.1 | — | 74.6 | 26 | |
| CLIPFeature=CLIP, Evaluation Protocol=Finetuned, Grounding head=CG-DETR2024.03 | 52.77 | 45.85 | 65.62 | 30.16 | |
| VDIYear=2023, Setup=FS2024.03 | 52.32 | — | — | 31.37 | |
| Huang et al.Year=2023, Setup=WS2024.03 | 52.18 | 45.2 | 69.16 | 23.94 | |
| Moment-DETRType=VLP, Size=N/A, FT=true, extra_pretraining=true2025.08 | 52.1 | — | 65.8 | 30.6 | |
| Moment-DETRYear=2021, Setup=FS2024.03 | 52.07 | 45.54 | 65.83 | 30.59 | |
| Moment-DETR2023.11 | 52.07 | 45.54 | 65.83 | 30.59 | |
| DEViL2025.12 | 51.5 | 47.7 | 72.6 | 25.2 | |
| CPLYear=2022, Setup=WS2024.03 | 49.05 | 43.23 | 65.99 | 22.61 | |
| VideoChat-TLLM Size=7B, Training Mode=Zero-shot2024.10 | 48.7 | — | 69.9 | 24 | |
| TimeChatFine-tuned=true2024.03 | 46.7 | — | — | 23.7 | |
| TimechatLLM Size=7B, Training Mode=Fine-tuned2024.10 | 46.7 | — | — | 23.7 | |
| TimeChatModel Category=Video-LLMs2025.10 | 46.7 | — | — | 23.7 | |
| 2D-TAN2023.11 | 46.02 | 41.25 | 58.76 | 27.4 | |
| ET-Chat# Train=164K, Protocol=zero-shot2026.06 | 45.9 | 42.3 | 65.7 | 20 | |
| 2D-TANType=VLP, Size=N/A, FT=true, extra_pretraining=true2025.08 | 45.8 | — | 57.3 | 27.9 | |
| 2D-TANYear=2020, Setup=FS2024.03 | 45.75 | 41.05 | 57.31 | 27.88 | |
| LLaVA-STLLM Scale=7B2025.01 | 44.8 | 42.4 | 63.1 | 23.4 | |
| E.M.GroundEvaluation Protocol=Zero-shot, Model Size=3.8B2026.02 | 44.8 | 43.3 | 67.1 | 19.4 | |
| LLaVA-ST2025.12 | 44.8 | 42.4 | 63.1 | 23.4 | |
| DRIFT (ET-Chat)# Train=100K, Protocol=zero-shot2026.06 | 44.8 | 43.8 | 67.2 | 20.6 | |
| VTG-GPTYear=2023, Setup=ZS2024.03 | 43.68 | 39.81 | 59.48 | 25.94 | |
| Luo et al.Year=2023, Setup=ZS2024.03 | 42.93 | 37.92 | 56.77 | 20.13 | |
| VSLNet2023.11 | 42.69 | 41.58 | 60.3 | 24.14 | |
| ARROWGEVModel Size=3B2026.01 | 42.5 | 42.3 | 64.9 | 19.7 | |
| LongVA-7B-DPOModel Category=General Vid-LLMs, Enhancement=NumPro-FT2024.11 | 42 | 41.4 | 63.8 | 20.6 | |
| TRACEEvaluation Protocol=Zero-shot, Model Size=7B2026.02 | 40.3 | — | — | 19.4 | |
| TRACEType=SFT, Size=7B, FT=false, extra_pretraining=false2025.08 | 40.3 | — | — | 19.4 | |
| Time-R1Model Size=3B2026.01 | 40 | 40.7 | 62.6 | 18.2 | |
| ED-VTG# Train=136K, Protocol=zero-shot2026.06 | 39.3 | 40.2 | 59.5 | 19.8 | |
| Kim et al.Year=2023, Setup=US2024.03 | 37.24 | 36.05 | 52.95 | 19.33 | |
| LT-ZVGProtocol=zero-shot2026.06 | 37.2 | 36 | 52.9 | 19.3 | |
| Qwen2-VL-7BModel Category=General Vid-LLMs, Enhancement=NumPro2024.11 | 36.8 | 38.5 | 60.7 | 15.9 | |
| Grounded-VideoLLMLLM Scale=4B2025.01 | 36.4 | 36.8 | 54.2 | 19.7 | |
| Grounded-VideoLLM2025.12 | 36.4 | 36.8 | 54.2 | 19.7 | |
| GPT-4oModel Category=General Vid-LLMs, Enhancement=NumPro2024.11 | 35.5 | 37.6 | 57.1 | 13.5 | |
| CNMYear=2022, Setup=WS2024.03 | 35.15 | 38.11 | 60.04 | 14.95 | |
| VTimeLLMEvaluation Protocol=Zero-shot, Model Size=13B2026.02 | 34.3 | 34.6 | 55.3 | 14.7 | |
| VTG-LLMModel Category=VTG-Tuned Vid-LLMs2024.11 | 33.8 | — | 52 | 15.7 | |
| VTG-LLMLLM Scale=7B2025.01 | 33.8 | — | — | 15.7 | |
| VTG-LLMEvaluation Protocol=Zero-shot, Model Size=7B2026.02 | 33.8 | — | — | 15.7 | |
| VTG-LLM2025.12 | 33.8 | — | — | 15.7 | |
| PZVMRYear=2022, Setup=US2024.03 | 33.21 | 32.62 | 46.83 | 18.51 | |
| ChatVTGLLM Size=7B, Training Mode=Zero-shot2024.10 | 33 | — | 52.7 | 15.9 | |
| ChatVTGEvaluation Protocol=Zero-shot, Model Size=7B2026.02 | 33 | 34.9 | 52.7 | 15.9 | |
| ChatVTG# Train=100K, Protocol=zero-shot2026.06 | 33 | 34.9 | 52.7 | 15.9 | |
| TimeChatZero-shot=true2024.03 | 32.2 | — | — | 13.4 | |
| TimeChatLLM Size=7B, Training Mode=Zero-shot2024.10 | 32.2 | — | — | 13.4 | |
| TimeChatLLM Scale=7B2025.01 | 32.2 | — | — | 13.4 | |
| TimeChatEvaluation Protocol=Zero-shot, Model Size=7B2026.02 | 32.2 | — | — | 13.4 | |
| TimeChat2025.12 | 32.2 | — | — | 13.4 | |
| TimeChat# Train=125K, Protocol=zero-shot2026.06 | 32.2 | — | — | 13.4 | |
| GPT-4oModel Category=General Vid-LLMs2024.11 | 32 | 35.4 | 55 | 11.5 | |
| HawkEyeZero-shot=true2024.03 | 31.4 | 33.7 | 50.6 | 14.5 | |
| HawkEyeLLM Size=7B, Training Mode=Zero-shot2024.10 | 31.4 | — | 50.6 | 14.5 | |
| HawkEyeModel Category=VTG-Tuned Vid-LLMs2024.11 | 31.4 | 33.7 | 50.6 | 14.5 | |
| HawkEyeLLM Scale=7B2025.01 | 31.4 | 33.7 | 50.6 | 14.5 | |
| HawkEyeEvaluation Protocol=Zero-shot, Model Size=7B2026.02 | 31.4 | 33.7 | 50.6 | 14.5 | |
| HawkEye2025.12 | 31.4 | 33.7 | 50.6 | 14.5 | |
| HawkEye# Train=715K, Protocol=zero-shot2026.06 | 31.4 | 33.7 | 50.6 | 14.5 | |
| PSVLProtocol=zero-shot2026.06 | 31.3 | 31.2 | 46.2 | 14.2 | |
| PSVLYear=2021, Setup=US2024.03 | 31.29 | 31.24 | 46.47 | 14.17 | |
| Seq2TimeEvaluation Protocol=Zero-shot, Model Size=7B2026.02 | 31.2 | — | — | 13.7 | |
| GroundingGPTLLM Size=7B, Training Mode=Zero-shot2024.10 | 29.6 | — | — | 11.9 | |
| GroundingGPTModel Category=VTG-Tuned Vid-LLMs2024.11 | 29.6 | — | — | 11.9 | |
| GroundingGPTLLM Scale=7B2025.01 | 29.6 | 28.7 | — | 11.9 | |
| DSCNetYear=2022, Setup=US2024.03 | 28.73 | — | 44.15 | 14.67 | |
| VTimeLLM-7BZero-shot=true2024.03 | 27.5 | 31.2 | 51 | 11.4 | |
| VTimeLLMLLM Size=7B, Training Mode=Zero-shot2024.10 | 27.5 | — | 51 | 11.4 |