Video Multimodal Understanding on MMVU
68.6AccuracySDRL
Evaluation Results
| Method | Links | |
|---|---|---|
| SDRLTraining=RL2026.03 | 68.6 | |
| VideoRFTTraining=SFT+ RL, Input frames=16-frame2026.03 | 67.3 | |
| TW-GRPOTraining=RL2026.03 | 65.8 | |
| Qwen2.5-VL-7BTraining=None2026.03 | 65.4 | |
| SDRLTraining=RL, Training dataset=EventFlow2026.03 | 64.8 | |
| Video-R1Training=SFT+ RL2026.03 | 64.2 | |
| VideoChat-R1Training=RL2026.03 | 64.2 | |
| Video-R1Training=RL2026.03 | 63.8 | |
| VideoRFTTraining=RL, Input frames=16-frame2026.03 | 63.5 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=ours CoT2026.03 | 63.2 | |
| VideoRFTTraining=SFT, Input frames=16-frame2026.03 | 60.5 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=video-r1 CoT2026.03 | 59.2 | |
| Video-R1Training=SFT2026.03 | 51.3 | |
| LLaVA-OneVision-7BTraining=None2026.03 | 49.2 | |
| MSA-PTTraining Strategy=from-scratch sparse pretraining2026.06 | 47.5 | |
| FullTraining Strategy=Full-Attention baseline2026.06 | 45.8 | |
| MSA-CPTTraining Strategy=sparse continued pretraining2026.06 | 45.8 | |
| VideoLLaMA2Training=None2026.03 | 44.8 |