Grounded Video Question Answering on NExT-GQA
61.2mIoUHuman
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HumanSupervision=Natural performance2024.09 | 61.2 | — | 70.3 | — | 86.2 | 72.1 | 82 | 93 | — | — | — | |
| Human2026.06 | 61.2 | — | — | — | 86.2 | 72.1 | 82.1 | — | — | — | — | |
| DynFrame-8BSize=8B2026.05 | 44.3 | — | — | — | — | — | — | — | — | — | 80 | |
| VITALSize=7B2026.05 | 43 | — | — | — | — | — | — | — | — | — | 78.7 | |
| DynFrame-4BSize=4B2026.05 | 41.5 | — | — | — | — | — | — | — | — | — | 77.6 | |
| Temporal-RLTSize=7B2026.05 | 37.3 | — | — | — | — | — | — | — | — | — | 78.7 | |
| DeepVideo-R1-7BModel parameters=7B2026.02 | 36.8 | — | — | — | — | — | 72.5 | — | — | — | — | |
| Qwen3-VL (Thinking)Size=8B2026.05 | 35.1 | — | — | — | — | — | — | — | — | — | 75.4 | |
| VideoTemp-o3-7B-RLModel parameters=7B, Training stage=RL2026.02 | 33.4 | — | — | — | — | — | 76.4 | — | — | — | — | |
| Qwen3-VL (Thinking)Size=4B2026.05 | 33.4 | — | — | — | — | — | — | — | — | — | 73.8 | |
| APPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 32.9 | — | — | — | — | — | — | 75.8 | — | — | — | |
| VideoChat-R1-7BModel parameters=7B2026.02 | 32.4 | — | — | — | — | — | 70.6 | — | — | — | — | |
| DAPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 32.4 | — | — | — | — | — | — | 75 | — | — | — | |
| GRPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 32 | — | — | — | — | — | — | 75.1 | — | — | — | |
| VideoMindSize=7B2025.03 | 31.4 | 50.2 | 25.8 | 56 | 35.3 | 39 | 28.2 | — | — | — | — | |
| Qwen2.5-VLSize=7B2026.05 | 30.5 | — | — | — | — | — | — | — | — | — | 76.5 | |
| LOVE-R1Size=7B2026.05 | 30.5 | — | — | — | — | — | — | — | — | — | 73 | |
| VideoTemp-o3-7B-SFTModel parameters=7B, Training stage=SFT2026.02 | 30.3 | — | — | — | — | — | 75.4 | — | — | — | — | |
| InternVL3Size=8B2026.05 | 30 | — | — | — | — | — | — | — | — | — | 80.4 | |
| VideoMindSize=2B2025.03 | 28.6 | 45.2 | 23.2 | 51.3 | 32.6 | 36.4 | 25.2 | — | — | — | — | |
| SFTBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 28.4 | — | — | — | — | — | — | 72.1 | — | — | — | |
| VideoChat-TPOSize=7B2025.03 | 27.7 | 41.2 | 23.4 | 51.3 | 32.8 | 35.6 | 25.5 | — | — | — | — | |
| HawkEyeSize=7B2025.03 | 25.7 | 37 | 19.5 | — | — | — | — | — | — | — | — | |
| TOGA2026.06 | 24.4 | — | — | — | 40.6 | 40.5 | 24.6 | — | — | — | — | |
| DeVi-Gemini-2.0Supervision=Zero-shot, Backbone=Gemini-2.02024.09 | 23.6 | — | 19.5 | — | 38.9 | 39.7 | 28.9 | 73.1 | — | — | — | |
| Qwen2.5-VL-7BModel parameters=7B2026.02 | 22.7 | — | — | — | — | — | 74.8 | — | — | — | — | |
| DeVi-GPT-4oSupervision=Zero-shot, Backbone=GPT-4o2024.09 | 22.3 | — | 17.4 | — | 37.9 | 39.3 | 28 | 71.6 | — | — | — | |
| VideoChat-R1Size=7B, Training Data=18K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 22 | — | — | — | — | — | — | 63.2 | — | — | — | |
| SeViLASize=4B2025.03 | 21.7 | 29.2 | 13.8 | 34.7 | 22.9 | 29.5 | 16.6 | — | — | — | — | |
| SeViLASupervision=Weakly-supervised, pre-trained on video grounding=true2024.09 | 21.7 | — | 13.8 | — | 22.9 | 29.5 | 16.6 | 68.1 | — | — | — | |
| SeViLA†2026.06 | 21.7 | — | — | — | 22.9 | 29.5 | 16.6 | — | — | — | — | |
| LLoViSupervision=Zero-shot2024.09 | 21.5 | — | 16.2 | — | 38 | 39.4 | 26.8 | 73.8 | — | — | — | |
| Random2026.06 | 21.1 | — | — | — | 8.7 | 21.1 | 1.7 | — | — | — | — | |
| Grounded-VideoLLM†‡2026.06 | 21.1 | — | — | — | 34.4 | 34.5 | 26.7 | — | — | — | — | |
| LeAdQA-7B‡Backbone=7B2026.06 | 20.5 | — | — | — | 29.5 | 30.3 | 19.2 | — | — | — | — | |
| CREDiT2026.06 | 20.1 | — | — | — | 41.1 | 41.4 | 27.9 | — | — | — | — | |
| LLoViSize=1.8T2025.03 | 20 | — | 15.3 | — | 36.9 | 37.3 | 24.3 | — | — | — | — | |
| LLoVi2026.06 | 20 | — | — | — | 36.9 | 37.3 | 24.3 | — | — | — | — | |
| VideoStreamingSize=8.3B2025.03 | 19.3 | — | 13.3 | — | 31 | 32.2 | 17.8 | — | — | — | — | |
| VideoStreamingSupervision=Zero-shot2024.09 | 19.3 | — | 13.3 | — | 31 | 32.2 | 17.8 | — | — | — | — | |
| VideoStreaming†2026.06 | 19.3 | — | — | — | 31 | 32.2 | 17.8 | — | — | — | — | |
| LangRepoSize=8x7B2025.03 | 18.5 | — | 12.2 | — | 28.7 | 31.3 | 17.1 | — | — | — | — | |
| GRPO-CARESize=7B, Training Data=260K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 18 | — | — | — | — | — | — | 51.5 | — | — | — | |
| TW-GRPOSize=7B, Training Data=1K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 17.7 | — | — | — | — | — | — | 53.2 | — | — | — | |
| Video-R1-7BModel parameters=7B2026.02 | 17.5 | — | — | — | — | — | 74.3 | — | — | — | — | |
| LongVTSize=7B2026.05 | 17.4 | — | — | — | — | — | — | — | — | — | 70.4 | |
| APPOSize=7B, Training Data=34K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 16.9 | — | — | — | — | — | — | 76.3 | — | — | — | |
| QGAC-TRSupervision=Weakly-supervised2024.09 | 15.7 | — | 11.7 | — | 27.7 | 28.3 | 18.3 | 63.6 | — | — | — | |
| Video-R1Size=7B, Training Data=260K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 14.9 | — | — | — | — | — | — | 64.7 | — | — | — | |
| VideoRFTSize=7B, Training Data=310K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 14.4 | — | — | — | — | — | — | 72 | — | — | — | |
| IGVSupervision=Weakly-supervised2024.09 | 14 | — | 9.6 | — | 18.9 | 21.4 | 10.2 | 50.1 | — | — | — | |
| IGV2026.06 | 14 | — | — | — | 18.9 | 21.4 | 10.2 | — | — | — | — | |
| Base ModelBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 13.5 | — | — | — | — | — | — | 49.9 | — | — | — | |
| CRAFrozenBiLM2026.06 | 13.5 | — | — | — | 25.9 | 26.5 | 18.8 | — | — | — | — | |
| FrozenBiLM(TimeCraft)Supervision=Weakly-supervised2024.09 | 13.2 | — | 8.4 | — | 24.9 | 26.3 | 18.5 | 74.7 | — | — | — | |
| TimeCraftFrozenBiLM2026.06 | 13.2 | — | — | — | 24.9 | 26.3 | 18.5 | — | — | — | — | |
| Temp[CLIP](NG+)Supervision=Weakly-supervised2024.09 | 12.6 | — | 8.9 | — | 25.5 | 25.7 | 15.9 | 60.2 | — | — | — | |
| APPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 11.1 | — | — | — | — | — | — | 71.2 | — | — | — | |
| DAPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 10.6 | — | — | — | — | — | — | 70.9 | — | — | — | |
| SFTBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 10.5 | — | — | — | — | — | — | 63.7 | — | — | — | |
| GRPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 10.3 | — | — | — | — | — | — | 70.7 | — | — | — | |
| Base ModelBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 10.1 | — | — | — | — | — | — | 40.3 | — | — | — | |
| FrozenBiLM NG+Size=890M2025.03 | 9.6 | 13.5 | 6.1 | 28.5 | 23.7 | 24.2 | 17.5 | — | — | — | — | |
| NG+FrozenBiLM2026.06 | 9.6 | — | — | — | 23.7 | 24.2 | 17.5 | — | — | — | — | |
| FrozenBiLM(NG+)Supervision=Weakly-supervised2024.09 | 9.5 | — | 6.1 | — | 23.7 | 24.2 | 17.5 | 70.8 | — | — | — | |
| VIOLETv22025.03 | 3.1 | 4.3 | 1.3 | 25.1 | 23.3 | 23.8 | 12.8 | — | — | — | — | |
| VIOLETv22026.06 | 3.1 | — | — | — | 23.3 | 23.6 | 12.8 | — | — | — | — | |
| VGTSupervision=Weakly-supervised2024.09 | 3 | — | 1.7 | — | 25.3 | 25.3 | 14.4 | 55.7 | — | — | — | |
| VGT[RBT]2026.06 | 3 | — | — | — | 25.3 | 25.3 | 14.4 | — | — | — | — | |
| Base ModelBackbone=Qwen3-VL-4B-Instruct2026.05 | — | — | — | — | — | — | — | — | 23.09 | 15.4 | — | |
| Base ModelBackbone=Qwen3-VL-8B-Instruct2026.05 | — | — | — | — | — | — | — | — | 35.93 | 21.83 | — | |
| V-ZeroBackbone=Qwen3-VL-4B-Instruct2026.05 | — | — | — | — | — | — | — | — | 22.82 | 15.22 | — | |
| V-ZeroBackbone=Qwen3-VL-8B-Instruct2026.05 | — | — | — | — | — | — | — | — | 35.73 | 22.11 | — | |
| Video-R1-7BSize=7B2026.05 | — | — | — | — | — | — | — | — | — | — | 77.3 | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=12026.05 | — | — | — | — | — | — | — | — | 31.66 | 20.8 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=22026.05 | — | — | — | — | — | — | — | — | 36.59 | 22.8 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=32026.05 | — | — | — | — | — | — | — | — | 38.14 | 23.25 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=42026.05 | — | — | — | — | — | — | — | — | 38.29 | 23 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=52026.05 | — | — | — | — | — | — | — | — | 37.76 | 22.82 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=12026.05 | — | — | — | — | — | — | — | — | 37.98 | 23.55 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=22026.05 | — | — | — | — | — | — | — | — | 39.04 | 24.31 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=32026.05 | — | — | — | — | — | — | — | — | 39.29 | 24.24 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=42026.05 | — | — | — | — | — | — | — | — | 39.6 | 24.26 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=52026.05 | — | — | — | — | — | — | — | — | 39.96 | 25.1 | — | |
| VideoChat-R1.5Size=7B2026.05 | — | — | — | — | — | — | — | — | — | — | 79.9 | |
| VideoChat-R1.5†‡2026.06 | — | — | — | — | 77.9 | 74.9 | 61.9 | — | — | — | — | |
| VisPlayBackbone=Qwen3-VL-4B-Instruct2026.05 | — | — | — | — | — | — | — | — | 23.14 | 15.33 | — | |
| VisPlayBackbone=Qwen3-VL-8B-Instruct2026.05 | — | — | — | — | — | — | — | — | 35.87 | 22.19 | — |