Video Reasoning on MMVU mc
82.6ScoreGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Size=-, # Frames=-2026.01 | 82.6 | |
| Gemini-2.5-ProSize=-, # Frames=-2026.01 | 78.4 | |
| GPT-4oSize=-, # Frames=-2026.01 | 75.4 | |
| Gemini-1.5-ProSize=-, # Frames=-2026.01 | 71.2 | |
| ProcessThinker (process-only)Backbone=QWEN3-VL-8B, Training strategy=GRPO, Reward type=process-only2026.04 | 68.48 | |
| ProcessThinker (outcome + process)Backbone=QWEN3-VL-8B, Training strategy=GRPO, Reward type=outcome + process2026.04 | 67.52 | |
| ProcessThinker (outcome-only)Backbone=QWEN3-VL-8B, Training strategy=GRPO, Reward type=outcome-only2026.04 | 67.36 | |
| Video-KTRSize=7B, # Frames=642026.01 | 66.6 | |
| Video-RTSSize=7B, # Frames=51.22026.01 | 66.4 | |
| VIDEO-R1-7BBackbone=QWEN2.5-VL2026.04 | 65.92 | |
| Video-KTRSize=7B, # Frames=322026.01 | 65.9 | |
| TW-GRPOSize=7B, # Frames=162026.01 | 65.8 | |
| Video-KTRSize=7B, # Frames=162026.01 | 65.7 | |
| QWEN3-VL-8B-INSTRUCTBackbone=QWEN3-VL-8B2026.04 | 65.6 | |
| OursSize=7B, #Frames=322026.06 | 65.3 | |
| PROCESSTHINKER-SFTBackbone=QWEN3-VL-8B, Training strategy=SFT2026.04 | 64.48 | |
| VideoChat-R1Size=7B, #Frames=162026.06 | 64.2 | |
| Video-R1Size=7B, # Frames=322026.01 | 63.8 | |
| Video-R1Size=7B, #Frames=322026.06 | 63.8 | |
| Ours(wo-mv)Size=7B, #Frames=322026.06 | 63.8 | |
| Qwen2.5-VL-SFTSize=7B, # Frames=322026.01 | 63.5 | |
| Qwen2.5-VL-7B-SFTSize=7B, #Frames=322026.06 | 63.5 | |
| Video-R1-wo-imageSize=7B, #Frames=322026.06 | 62.7 | |
| Qwen2.5-VL-SFTSize=7B, # Frames=642026.01 | 61.6 | |
| Qwen2.5-VL-SFTSize=7B, # Frames=162026.01 | 61.3 | |
| Video-R1-wo-imageSize=7B, #Frames=162026.06 | 60.6 | |
| Qwen2.5-VLSize=7B, # Frames=-2026.01 | 59.2 | |
| Qwen2.5-VL-7BSize=7B2026.06 | 59.2 | |
| LLaVA-OVSize=7B, # Frames=642026.01 | 49.2 | |
| VILA-1.5Size=8B, # Frames=642026.01 | 49.2 | |
| LLaVA-OneVisionSize=7B, #Frames=642026.06 | 49.2 | |
| VideoLLaMA2Size=8×7B, #Frames=162026.06 | 44.8 |