Video Question Answering on NExT-GQA
79.3AccuracySDRL
Evaluation Results
| Method | Links | |
|---|---|---|
| SDRLTraining=RL2026.03 | 79.3 | |
| SDRLTraining=RL, Training dataset=EventFlow2026.03 | 77.3 | |
| TW-GRPOTraining=RL2026.03 | 76.1 | |
| VideoChat-R1Training=RL2026.03 | 76 | |
| Qwen2.5-VL-7BTraining=None2026.03 | 75.9 | |
| VideoRFTTraining=SFT+ RL, Input frames=16-frame2026.03 | 75.1 | |
| Video-R1Training=SFT+ RL2026.03 | 74.3 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=ours CoT2026.03 | 73.6 | |
| Factum-4BSource Type=Open-Source, FPS=1fps2026.04 | 73.6 | |
| Qwen3-VL-4B-InstructSource Type=Open-Source, FPS=1fps2026.04 | 72.1 | |
| Qwen3-VL-4B-ThinkingSource Type=Open-Source, FPS=1fps2026.04 | 66.6 | |
| Qwen2.5-VL-7BSource Type=Open-Source, FPS=1fps2026.04 | 59.5 | |
| VTimeLLM-7BSource Type=Open-Source, FPS=1fps2026.04 | 17.4 |