Video Reasoning on Video-MMMU
84.6AccuracyGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Size=-, # Frames=-2026.01 | 84.6 | — | |
| Gemini-2.5-ProSize=-, # Frames=-2026.01 | 83.6 | — | |
| Seed1.5-VLThinking mode=true, Evaluation frame rate=1 FPS2025.05 | 81.4 | — | |
| Kimi-K1.6Evaluation frame rate=1 FPS2025.05 | 76.7 | — | |
| GLM-4.5V2026.04 | 72.4 | — | |
| Seed1.5-VLThinking mode=false, Evaluation frame rate=1 FPS2025.05 | 72.1 | — | |
| EFlow2026.07 | 65.5 | — | |
| Qwen3-VL-8BModel size=8B2026.07 | 65.3 | — | |
| ProcessThinker (process-only)Backbone=QWEN3-VL-8B, Training strategy=GRPO, Reward type=process-only2026.04 | 63.33 | — | |
| QWEN3-VL-8B-INSTRUCTBackbone=QWEN3-VL-8B2026.04 | 62.89 | — | |
| ProcessThinker (outcome + process)Backbone=QWEN3-VL-8B, Training strategy=GRPO, Reward type=outcome + process2026.04 | 61.67 | — | |
| GPT-4oSize=-, # Frames=-2026.01 | 61.2 | — | |
| GPT-4oModel Category=Closed-Source MLLMs2026.01 | 61.2 | — | |
| GPT-4o2026.04 | 61.2 | — | |
| GPT-4o2026.04 | 61.2 | — | |
| ProcessThinker (outcome-only)Backbone=QWEN3-VL-8B, Training strategy=GRPO, Reward type=outcome-only2026.04 | 60.78 | — | |
| PROCESSTHINKER-SFTBackbone=QWEN3-VL-8B, Training strategy=SFT2026.04 | 58.78 | — | |
| AdaFocusReasoning Mode=AutoThink2026.05 | 56.8 | — | |
| FFRDistillation Method=FFR, Teacher Model=Qwen3-235B2026.04 | 56.5 | — | |
| VideoAuto-R1Reasoning Mode=AutoThink2026.05 | 55.6 | — | |
| FFRBase Model=VideoRFT-SFT, Training=RFT + FFR2026.04 | 54.9 | 13.2 | |
| Qwen2.5-VL-7B (Base)Reasoning Mode=✗2026.05 | 54.7 | — | |
| FFRBase Model=Video-R1-SFT, Training=RFT + FFR2026.04 | 54.6 | 15.19 | |
| FFR (GLM-4.5V)Distillation Method=FFR, Teacher Model=GLM-4.5V2026.04 | 54.6 | — | |
| RLER2026.04 | 54.2 | — | |
| VITAL-7BModel size=7B2026.07 | 54.2 | — | |
| Gemini-1.5-ProModel Category=Closed-Source MLLMs2026.01 | 53.9 | — | |
| VIDEO-R1-7BBackbone=QWEN2.5-VL2026.04 | 53.89 | — | |
| Gemini-1.5-ProSize=-, # Frames=-2026.01 | 53.4 | — | |
| Video-KTRSize=7B, # Frames=642026.01 | 53.1 | — | |
| Video-RTSSize=7B, # Frames=51.22026.01 | 52.7 | — | |
| Video-KTRSize=7B, # Frames=322026.01 | 52.6 | — | |
| Video-R1-7BModel size=7B2026.07 | 52.4 | — | |
| Video-R1Size=7B, # Frames=322026.01 | 52.3 | — | |
| Video-R12026.04 | 52.3 | — | |
| Video-R12026.04 | 52.3 | — | |
| DPS (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DPS2026.01 | 51.6 | — | |
| VideoChat-R1.52026.04 | 51.4 | — | |
| Video-R1Reasoning Mode=Think-Only2026.05 | 51.4 | — | |
| TW-GRPOSize=7B, # Frames=162026.01 | 51.3 | — | |
| Video-KTRSize=7B, # Frames=162026.01 | 51.3 | — | |
| VideoRFTModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 51.1 | — | |
| VideoRFT2026.04 | 51.1 | — | |
| VideoRFT2026.04 | 51.1 | — | |
| VideoRFT-7BModel size=7B2026.07 | 51.1 | — | |
| FFRDistillation Method=FFR, Teacher Model=Qwen3-32B2026.04 | 50.5 | — | |
| MOSS-ChatV2026.04 | 50.2 | — | |
| Two-stage RL (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=Two-stage RL (DPS and annealing)2026.01 | 50.1 | — | |
| VideoChat-R12026.04 | 50 | — | |
| VideoR1Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 49.8 | — | |
| VideoChat-R1.5Reasoning Mode=Think-Only2026.05 | 49.6 | — | |
| Qwen2.5-VL-SFTSize=7B, # Frames=322026.01 | 49.4 | — | |
| Qwen2.5-VL-SFTSize=7B, # Frames=642026.01 | 49.4 | — | |
| Pixel-Reasoner2026.04 | 49.3 | — | |
| STAR-R12026.04 | 49.2 | — | |
| DAPO (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DAPO2026.01 | 49 | — | |
| Video-ChatR12026.04 | 48.9 | — | |
| VideoRFT-SFT2026.04 | 48.5 | — | |
| Qwen2.5-VL2026.04 | 47.8 | — | |
| Qwen2.5-VLSize=7B, # Frames=-2026.01 | 47.4 | — | |
| Qwen2.5-VL-SFTSize=7B, # Frames=162026.01 | 47.4 | — | |
| Qwen2.5-VL-7B2026.04 | 47.4 | — | |
| Video-R1-SFT2026.04 | 47.4 | — | |
| Video-R1-SFTDistillation Method=SFT, Teacher Model=Baseline2026.04 | 47.4 | — | |
| Qwen2.5-VL-7BModel size=7B2026.07 | 47.4 | — | |
| SFTDistillation Method=SFT, Teacher Model=Qwen3-235B2026.04 | 46.2 | — | |
| VideoLLaMA 3-7B2026.04 | 46 | — | |
| Qwen2.5-VLModel Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 45.8 | — | |
| Video-Thinker2026.04 | 43.9 | — | |
| SFTDistillation Method=SFT, Teacher Model=Qwen3-32B2026.04 | 42 | — | |
| TW-GRPOModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 40.8 | — | |
| LLaVA-Video-7B2026.04 | 36.1 | — | |
| InternVL2.5Model Category=Open-Source General MLLMs, Parameter Scale=8B2026.01 | 35.2 | — | |
| LLaVA-OVSize=7B, # Frames=642026.01 | 33.8 | — | |
| VILA-1.5Size=8B, # Frames=642026.01 | 33.8 | — | |
| LLaVA-OneVision-7B2026.04 | 33.8 | — | |
| LLaVA-OV-7BModel size=7B2026.07 | 33.8 | — | |
| mPLUG-Owl3Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 32 | — | |
| LongVA-7BModel size=7B2026.07 | 23.9 | — | |
| LLaVA-NeXT-InterleaveModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 23.2 | — | |
| VILA-1.5-8B2026.04 | 20.9 | — | |
| VILA-1.5-8BModel size=8B2026.07 | 20.8 | — | |
| Mantis-Idefics2Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 19.3 | — |