Temporal Video Grounding on Charades-STA
70.3Rank-1 Recall (IoU=0.5)Bridge-STG
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Bridge-STGModel Architecture Type=7B-based MLLM, Temporal Modality=Spatio-Temporal2026.04 | 70.3 | — | 49.3 | |
| MeCoEvaluation Protocol=Dataset-wise fine-tuning, Backbone=QWen2VL, Model Scale=7B2025.03 | 68.5 | 82.3 | 41.6 | |
| EaTRModel Architecture Type=Non-generative and task-specific, Temporal Modality=only-Temporal2026.04 | 68.4 | — | 44.9 | |
| VideoChat-TEvaluation Protocol=Dataset-wise fine-tuning, Model Scale=7B2025.03 | 67.1 | 79.4 | 43 | |
| TimeSuiteModel Architecture Type=7B-based MLLM, Temporal Modality=only-Temporal2026.04 | 67.1 | — | 43 | |
| MeCoEvaluation Protocol=Dataset-wise fine-tuning, Backbone=ETChat, Model Scale=7B2025.03 | 63.9 | 77.2 | 40.1 | |
| SpaceVLLMModel Architecture Type=7B-based MLLM, Temporal Modality=Spatio-Temporal2026.04 | 63.6 | — | 38.5 | |
| TRACEEvaluation Protocol=Dataset-wise fine-tuning, Model Scale=7B2025.03 | 61.7 | — | 41.4 | |
| MeCoEvaluation Protocol=Dataset-wise fine-tuning, Backbone=ETChat, Model Scale=3.8B2025.03 | 61.6 | 75.3 | 38.5 | |
| UniVTGEvaluation Protocol=Dataset-wise fine-tuning2025.03 | 60.2 | 72.6 | 38.6 | |
| CG-DETREvaluation Protocol=Dataset-wise fine-tuning2025.03 | 58.4 | 70.4 | 36.3 | |
| HawkEyeEvaluation Protocol=Dataset-wise fine-tuning, Model Scale=7B2025.03 | 58.3 | 72.5 | 28.8 | |
| HawkEyeModel Architecture Type=7B-based MLLM, Temporal Modality=only-Temporal2026.04 | 58.3 | — | 28.8 | |
| QD-DETRModel Architecture Type=Non-generative and task-specific, Temporal Modality=only-Temporal2026.04 | 57.3 | — | 32.6 | |
| VTG-LLMEvaluation Protocol=Dataset-wise fine-tuning, Model Scale=7B2025.03 | 57.2 | — | 33.4 | |
| VTG-LLMModel Architecture Type=7B-based MLLM, Temporal Modality=only-Temporal2026.04 | 57.2 | — | 33.4 | |
| M-DETREvaluation Protocol=Dataset-wise fine-tuning2025.03 | 52.1 | 65.8 | 30.6 | |
| MeCoEvaluation Protocol=Zero-shot (Fine-tuned on E.T.Instruct), Backbone=QWen2VL, Model Scale=7B2025.03 | 50.1 | 71.1 | 23.3 | |
| VideoChat-TEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 48.7 | 69.9 | 24 | |
| TimeChatEvaluation Protocol=Dataset-wise fine-tuning, Model Scale=7B2025.03 | 46.7 | — | 23.7 | |
| MeCoEvaluation Protocol=Zero-shot (Fine-tuned on E.T.Instruct), Backbone=ETChat, Model Scale=7B2025.03 | 46.4 | 69.6 | 19.1 | |
| LLaVA-STModel Architecture Type=7B-based MLLM, Temporal Modality=Spatio-Temporal2026.04 | 44.8 | — | 23.4 | |
| MeCoEvaluation Protocol=Zero-shot (Fine-tuned on E.T.Instruct), Backbone=ETChat, Model Scale=3.8B2025.03 | 44.4 | 66.7 | 17.5 | |
| E.T.ChatEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=3.8B2025.03 | 43.2 | 64.4 | 19.4 | |
| E.T.ChatEvaluation Protocol=Zero-shot (Fine-tuned on E.T.Instruct), Model Scale=3.8B2025.03 | 43.2 | 64.4 | 19.4 | |
| NumPro-FTEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 42 | 63.8 | 20.6 | |
| TRACEEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 40.3 | — | 19.4 | |
| videollama (w. TCAS)Backbone=videollama, Evaluation Strategy=w. TCAS2025.10 | 37.5 | — | 20.1 | |
| videollamaBackbone=videollama, Evaluation Strategy=w. SFT2025.10 | 37.1 | — | 20.1 | |
| VTimeLLMEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=13B2025.03 | 34.3 | 55.3 | 14.7 | |
| VTG-LLMEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 33.8 | — | 15.7 | |
| Qwen2.5vl (w. TCAS)Backbone=Qwen2.5vl, Evaluation Strategy=w. TCAS2025.10 | 32.4 | — | 16.6 | |
| TimeChatEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 32.2 | — | 13.4 | |
| HawkEyeEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 31.4 | 50.6 | 4.5 | |
| Seq2TimeEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 31.2 | — | 13.7 | |
| GroundingGPTLLM Size=7B2024.01 | 29.6 | — | 11.9 | |
| VTimeLLMEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 27.5 | 51 | 11.4 | |
| Qwen2.5vlBackbone=Qwen2.5vl, Evaluation Strategy=w. SFT2025.10 | 27.3 | — | 14.1 | |
| MomentorEvaluation Protocol=Zero-shot (official checkpoints), Model Scale=7B2025.03 | 26.6 | 42.6 | 11.6 | |
| TimeChatEvaluation Protocol=Zero-shot (Fine-tuned on E.T.Instruct), Model Scale=7B2025.03 | 24.9 | 43.4 | 9.2 | |
| TRACEEvaluation Protocol=Zero-shot (Fine-tuned on E.T.Instruct), Model Scale=7B2025.03 | 23.7 | 39.4 | 11.5 | |
| CG-STVGModel Architecture Type=Non-generative and task-specific, Temporal Modality=Spatio-Temporal2026.04 | 20 | — | 7.1 | |
| TA-STVGModel Architecture Type=Non-generative and task-specific, Temporal Modality=Spatio-Temporal2026.04 | 16.3 | — | 5.2 | |
| VTGLLMEvaluation Protocol=Zero-shot (Fine-tuned on E.T.Instruct), Model Scale=7B2025.03 | 9.8 | 24.8 | 3.5 | |
| VideoChatGPTLLM Size=7B2024.01 | 7.7 | — | 1.7 | |
| Video-LLaMALLM Size=7B2024.01 | 3.8 | — | 0.9 | |
| VideoChatLLM Size=7B2024.01 | 3.3 | — | 1.3 |