Grounded Video Question Answering on NExT-GQA (test)
70.6Acc@GQAVideoChat-R1
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| VideoChat-R1Type=MLLM2025.12 | 70.6 | 36.1 | — | — | — | — | — | — | |
| TempR1Type=MLLM2025.12 | 70.1 | 39.2 | 58.9 | 35.9 | — | — | — | 17.1 | |
| MUSEGType=MLLM2025.12 | 63.5 | 25.2 | 34.7 | 16.4 | — | — | — | 7.4 | |
| VideoChat-R1.5Type=MLLM2025.12 | 62.4 | 20.5 | 28.2 | 12.7 | — | — | — | 4.6 | |
| Time-R1Type=MLLM2025.12 | 62.1 | 28.3 | 39.2 | 18.6 | — | — | — | 8.2 | |
| Qwen2.5-VL-7BType=MLLM2025.12 | 60.8 | 15.3 | 20.6 | 11.1 | — | — | — | 5.2 | |
| DEViL2025.12 | 36.3 | 27.9 | — | — | 48.1 | 48.1 | 23.4 | — | |
| TimeThinkSize=7B, Training Paradigm=Reinforcement Fine-Tuning, Zero-Shot=true2026.07 | 29.1 | 35.8 | — | — | 40.3 | — | — | — | |
| Zoom-ZeroModel Scale=7B-8B, Training Paradigm=RL-based2025.12 | 29 | 37.6 | 55.6 | 33.8 | — | — | — | — | |
| VideoMindSize=7B, Training Paradigm=Reinforcement Fine-Tuning, Zero-Shot=true2026.07 | 28.2 | 31.4 | — | — | 39 | — | — | — | |
| Qwen2.5-VL-GRPOSize=7B, Training Paradigm=Reinforcement Fine-Tuning, Zero-Shot=true2026.07 | 28.2 | 34 | — | — | 37.4 | — | — | — | |
| Qwen2.5-VL-SFTSize=7B, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 27.7 | 28.6 | — | — | 35.8 | — | — | — | |
| Grounded-VideoLLMModel Scale=7B-8B, Training Paradigm=SFT-based2025.12 | 26.7 | 21.1 | — | 18 | — | — | — | — | |
| Grounded-VideoLLM2025.12 | 26.7 | 21.1 | — | — | 34.5 | 34.4 | 18 | — | |
| Grounding-VideoLLMSize=4B, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 26.7 | 21.1 | — | — | 34.5 | — | — | — | |
| VideoChat-TPOModel Scale=7B-8B, Training Paradigm=RL-based2025.12 | 25.5 | 27.7 | 41.2 | 23.4 | — | — | — | — | |
| VideoChat-TPO2025.12 | 25.5 | 27.7 | — | — | 35.6 | 32.8 | 23.4 | — | |
| VideoChat-TPOSize=7B, Training Paradigm=Reinforcement Fine-Tuning, Zero-Shot=true2026.07 | 25.5 | 27.7 | — | — | 35.6 | — | — | — | |
| VideoMindSize=1.5B, Training Paradigm=Reinforcement Fine-Tuning, Zero-Shot=true2026.07 | 25.2 | 28.6 | — | — | 35.6 | — | — | — | |
| TOGA2026.04 | 24.6 | 24.4 | — | — | 40.5 | 40.6 | 21.1 | — | |
| TOGAmodel_type=classical expert model2026.04 | 24.6 | 24.4 | — | — | 40.5 | 40.6 | 21.1 | — | |
| VideoChat-R1Model Scale=7B-8B, Training Paradigm=RL-based2025.12 | 24.3 | 32.4 | 50.2 | 27.7 | — | — | — | — | |
| LLoVi2025.12 | 24.3 | 20 | — | — | 37.3 | 36.9 | 15.3 | — | |
| LLoViSize=1.8T, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 24.3 | — | — | — | 24.3 | — | — | — | |
| HawkEyeSize=7B, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 23.5 | 25.7 | — | — | 33.4 | — | — | — | |
| GroundVTS-QBackbone=Qwen2.5-VL-7B2026.04 | 23.2 | 25.8 | — | — | 37.4 | 35.4 | 20.4 | — | |
| GroundVTS-Qbackbone=Qwen2.5VL-7B2026.04 | 23.2 | 25.8 | — | — | 37.4 | 35.4 | 20.4 | — | |
| TVG-R1Model Scale=7B-8B, Training Paradigm=RL-based2025.12 | 22.1 | 29.2 | 41.6 | 20.8 | — | — | — | — | |
| Qwen2.5-VLModel Scale=7B-8B2025.12 | 18.9 | 20.2 | 31.6 | 18.1 | — | — | — | — | |
| GroundVTS-IBackbone=InternVL3.5-8B2026.04 | 18.5 | 16.7 | — | — | 26.5 | 24.3 | 11.9 | — | |
| GroundVTS-Ibackbone=InternVL3.5-8B2026.04 | 18.5 | 16.7 | — | — | 26.5 | 24.3 | 11.9 | — | |
| VideoStreamingSize=8.3B, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 18.1 | 19.4 | — | — | 30.8 | — | — | — | |
| VideoStreaming2025.12 | 17.8 | 19.3 | — | — | 32.2 | 31 | 13.3 | — | |
| VidStreaming2026.04 | 17.8 | 19.3 | — | — | 32.2 | 31 | 13.3 | — | |
| VideoStreaming2026.04 | 17.8 | 19.3 | — | — | 32.2 | 31 | 13.3 | — | |
| FrozenBiLM NG+2025.12 | 17.5 | 9.6 | — | — | 24.2 | 23.7 | 6.1 | — | |
| FrozenBiLM NG+Size=890M, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 17.5 | 9.6 | — | — | 24.2 | — | — | — | |
| VTimeLLMModel Scale=7B-8B, Training Paradigm=SFT-based2025.12 | 17.4 | 24.4 | 36.1 | 20.1 | — | — | — | — | |
| LangRepo2025.12 | 17.1 | 18.5 | — | — | 31.3 | 28.7 | 12.2 | — | |
| SeViLA2025.12 | 16.6 | 21.7 | — | — | 29.5 | 22.9 | 13.8 | — | |
| SeViLASize=4B, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 16.6 | 21.7 | — | — | 29.5 | — | — | — | |
| LangRepoSize=8×7B, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 14.9 | 18.5 | — | — | 27.1 | — | — | — | |
| VIOLETv22025.12 | 12.8 | 3.1 | — | — | 23.6 | 23.3 | 1.3 | — | |
| VIOLETv2Size=-, Training Paradigm=Supervised Fine-Tuning, Zero-Shot=true2026.07 | 12.8 | — | — | — | 23.6 | — | — | — | |
| TimeChatModel Scale=7B-8B, Training Paradigm=SFT-based2025.12 | 7.6 | 20.6 | 34.1 | 17.9 | — | — | — | — | |
| HawkEye2025.12 | — | 25.7 | — | — | — | — | 19.5 | — |