Video Understanding on MMVU
78.2AccuracySeed2.0 Pro
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Seed2.0 Pro2026.06 | 78.2 | — | — | — | — | |
| Gemini-3-Flash2026.06 | 77.9 | — | — | — | — | |
| Gemini-3-Pro2026.06 | 76.3 | — | — | — | — | |
| GPT-4o2026.05 | 75.4 | — | — | — | — | |
| Seed2.0 Lite2026.06 | 75 | — | — | — | — | |
| Qwen3-VL-8B + ADPOPost-training Strategy=ADPO2026.05 | 73.1 | — | — | — | — | |
| Seed1.82026.06 | 73.1 | — | — | — | — | |
| Qwen3-VL-8B + DAPOPost-training Strategy=DAPO2026.05 | 72 | — | — | — | — | |
| VideoSearcher-8BEvaluation Setting=Agentic Model2026.07 | 70.4 | — | — | — | — | |
| Qwen3-VL-8B + GRPOPost-training Strategy=GRPO2026.05 | 70.2 | — | — | — | — | |
| Qwen3-VL-8BPost-training Strategy=Base Model2026.05 | 69.7 | — | — | — | — | |
| Qwen3-VL-8B-InstructEvaluation Setting=Direct Answer2026.07 | 69.28 | — | — | — | — | |
| Seed2.0 Mini2026.06 | 69 | — | — | — | — | |
| VideoSearcher-4BEvaluation Setting=Agentic Model2026.07 | 68.64 | — | — | — | — | |
| ParaVT-8BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 68.6 | — | — | — | — | |
| VideoRFT-7B2026.05 | 68.5 | — | — | — | — | |
| VideoRFTEvaluation Setting=Direct Answer2026.07 | 68.5 | — | — | — | — | |
| Qwen3-VL-8BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 68 | — | — | — | — | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 22026.05 | 67 | — | — | — | — | |
| GPT-4oSetting=best-setting numbers from official reports2026.05 | 66.7 | — | — | — | — | |
| Gemini-2.0-Flash2026.05 | 66.5 | — | — | — | — | |
| Video-RTS-7B2026.05 | 66.4 | — | — | — | — | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 32026.05 | 66.2 | — | — | — | — | |
| ViSS-R1-7B2026.05 | 66.1 | — | — | — | — | |
| VidGroundFrames=322026.04 | 65.8 | — | — | — | — | |
| Gemini 1.5 ProSetting=best-setting reports2026.05 | 65.8 | — | — | — | — | |
| VidGroundFrames=642026.04 | 65.6 | — | — | — | — | |
| Qwen2.5-VL-7B + SAGESAGE=true2026.05 | 65.4 | — | — | — | — | |
| Qwen2.5-VL-7BSetting=direct-answer setting (native single-pass prompt)2026.05 | 65.4 | — | — | — | — | |
| VideoCoMEvaluation Setting=Direct Answer2026.07 | 65.4 | — | — | — | — | |
| Qwen2.5-VL-7B + SFT + SAGESFT=true, SAGE=true2026.05 | 65.3 | — | — | — | — | |
| Video-RTSFrames=322026.04 | 65 | — | — | — | — | |
| VideoChat-R1-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 65 | — | — | — | — | |
| Base ModelModel=Qwen3-VL-8B-Instruct, Training Iteration=w/o training2026.05 | 64.8 | — | — | — | — | |
| Video-Thinker-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 64.5 | — | — | — | — | |
| VidGroundFrames=162026.04 | 64.2 | — | — | — | — | |
| TW-GRPOFrames=642026.04 | 64.2 | — | — | — | — | |
| Video-R1-7B2026.05 | 64.2 | — | — | — | — | |
| Conan-7BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 64 | — | — | — | — | |
| Video-RTSFrames=642026.04 | 63.9 | — | — | — | — | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 12026.05 | 63.8 | — | — | — | — | |
| Video-R1Evaluation Setting=Direct Answer2026.07 | 63.8 | — | — | — | — | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Frozen Questioner2026.05 | 63.7 | — | — | — | — | |
| Time-R1-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 63.4 | — | — | — | — | |
| LongVT-RFT-7BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 63.4 | — | — | — | — | |
| Qwen3-VL-4B-InstructEvaluation Setting=Direct Answer2026.07 | 63.36 | — | — | — | — | |
| TW-GRPOFrames=322026.04 | 63.1 | — | — | — | — | |
| Qwen2.5-VL-7BFrames=642026.04 | 62.6 | — | — | — | — | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 32026.05 | 62.6 | — | — | — | — | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 32026.05 | 62.4 | — | — | — | — | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 22026.05 | 62.4 | — | — | — | — | |
| Qwen2.5-VL-7BFrames=322026.04 | 62.3 | — | — | — | — | |
| VideoChat-R1.5Evaluation Setting=Direct Answer2026.07 | 62.08 | — | — | — | — | |
| TW-GRPOFrames=162026.04 | 61.8 | — | — | — | — | |
| Video-RTSFrames=162026.04 | 61.8 | — | — | — | — | |
| VideoZoomer-7BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 61.6 | — | — | — | — | |
| LongVILA-R1-7BFrames=322026.04 | 61.5 | — | — | — | — | |
| Video-R1-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 61.3 | — | — | — | — | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Frozen Questioner2026.05 | 61.1 | — | — | — | — | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 12026.05 | 61 | — | — | — | — | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 12026.05 | 60.6 | — | — | — | — | |
| Qwen2.5-VL-7BFrames=162026.04 | 60.5 | — | — | — | — | |
| Qwen2.5-VL-7B + SFTSFT=true2026.05 | 60.5 | — | — | — | — | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Frozen Questioner2026.05 | 60.5 | — | — | — | — | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 22026.05 | 60.5 | — | — | — | — | |
| Base ModelModel=Qwen3-VL-4B-Instruct, Training Iteration=w/o training2026.05 | 60.3 | — | — | — | — | |
| ReWatch-R1-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 59.8 | — | — | — | — | |
| Qwen2.5-VL-7B2026.05 | 59.2 | — | — | — | — | |
| Base ModelModel=Qwen2.5-VL-7B-Instruct, Training Iteration=w/o training2026.05 | 59.2 | — | — | — | — | |
| LongVILA-R1-7BFrames=162026.04 | 59.1 | — | — | — | — | |
| LongVILA-R1-7BFrames=642026.04 | 58.8 | — | — | — | — | |
| Qwen3-VLModel Size=8B2026.03 | 58.7 | — | — | — | — | |
| Video-R1Frames=322026.04 | 56.2 | — | — | — | — | |
| SAGE-7BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 55.7 | — | — | — | — | |
| Qwen2.5-VL-7B-SFTFrames=642026.04 | 55.4 | — | — | — | — | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 32026.05 | 54.6 | — | — | — | — | |
| Video-R1Frames=162026.04 | 54.5 | — | — | — | — | |
| Penguin-VLModel Size=8B2026.03 | 53.9 | — | — | — | — | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Frozen Questioner2026.05 | 53.3 | — | — | — | — | |
| Video-R1Frames=642026.04 | 53.2 | — | — | — | — | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 22026.05 | 53.1 | — | — | — | — | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 12026.05 | 51.7 | — | — | — | — | |
| InternVL-3.5Model Size=8B2026.03 | 51.5 | — | — | — | — | |
| Qwen2.5-VL-7B-SFTFrames=162026.04 | 51.3 | — | — | — | — | |
| OpenAI GPT-5 nanoModel Size=nano2026.03 | 51 | — | — | — | — | |
| Qwen2.5-VL-7B-SFTFrames=322026.04 | 51 | — | — | — | — | |
| Human2026.06 | 49.7 | — | — | — | — | |
| LLaVA-OneVision-7B2026.05 | 49.2 | — | — | — | — | |
| Base ModelModel=Qwen2.5-VL-3B-Instruct, Training Iteration=w/o training2026.05 | 48.5 | — | — | — | — | |
| VideoLLaMA22026.05 | 44.8 | — | — | — | — | |
| VideoRFT-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 42.7 | — | — | — | — | |
| CapRL++Model Category=Specialized Video VLMs, Backbone=Qwen3-VL-4B2026.06 | — | — | — | — | 48.5 | |
| DELTAVID-Qwen3-VL-8BParameters=8B, Base Model=Qwen3-VL2026.06 | — | — | — | 65.4 | — | |
| Gemini 2.5 Pro2026.03 | — | — | — | 76.1 | — | |
| Gemini-3-Pro2026.03 | — | — | — | 76.3 | — | |
| LLaVA-NeXTBackbone=LLaVA-NeXT2026.01 | — | 31.3 | 30.5 | — | — | |
| LLaVA-NeXT + Dino-HealBackbone=LLaVA-NeXT2026.01 | — | 30.9 | 30.7 | — | — | |
| LLaVA-NeXT + MotionCDBackbone=LLaVA-NeXT2026.01 | — | 26.3 | 25.1 | — | — | |
| LLaVA-NeXT + TCDBackbone=LLaVA-NeXT2026.01 | — | 25.6 | 29.8 | — | — | |
| Qwen3-VL-235B-A22BModel Category=General VLMs, Variant=Instruct2026.06 | — | — | — | — | 51 |