Video Reasoning on Video-Holmes (Accuracy)
67.4AccuracySeed2.0 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Seed2.0 ProSubtitles=true2026.06 | 67.4 | |
| Gemini-3-FlashSubtitles=true2026.06 | 65.6 | |
| Seed1.8Subtitles=true2026.06 | 65.5 | |
| Gemini-3-ProSubtitles=true2026.06 | 64.2 | |
| Seed2.0 LiteSubtitles=true2026.06 | 63.8 | |
| VideoChat-M1Model Category=Ours, Number of Parameters=37B2025.11 | 60.5 | |
| Seed2.0 MiniSubtitles=true2026.06 | 58.6 | |
| OmniJigsaw (CMM)Inference Mode=w/ Audio2026.04 | 52.53 | |
| FFR (GLM-4.5V)Distillation Method=FFR, Teacher Model=GLM-4.5V2026.04 | 52.3 | |
| OmniJigsaw (SMS)Inference Mode=w/ Audio2026.04 | 52.26 | |
| VideoJigsawInference Mode=w/ Audio2026.04 | 51.99 | |
| FFRDistillation Method=FFR, Teacher Model=Qwen3-235B2026.04 | 51.6 | |
| Qwen3-Omni-30BInference Mode=w/ Audio2026.04 | 50.84 | |
| OmniJigsaw (JMI)Inference Mode=w/ Audio2026.04 | 50.24 | |
| OmniJigsaw (CMM)Inference Mode=w/o Audio2026.04 | 48.29 | |
| OmniJigsaw (SMS)Inference Mode=w/o Audio2026.04 | 47.96 | |
| VideoJigsawInference Mode=w/o Audio2026.04 | 47.9 | |
| FFRDistillation Method=FFR, Teacher Model=Qwen3-32B2026.04 | 47.8 | |
| OmniJigsaw (JMI)Inference Mode=w/o Audio2026.04 | 47.47 | |
| SFTDistillation Method=SFT, Teacher Model=Qwen3-235B2026.04 | 47.1 | |
| Qwen3-Omni-30BInference Mode=w/o Audio2026.04 | 46.6 | |
| Gemini 1.5 ProModel Category=Closed-source Large-sized MLLM2025.11 | 45.7 | |
| Gemini-1.5-ProFrames=322026.05 | 45.7 | |
| Mixed RLVRTraining Strategy=Mixed RLVR2026.04 | 45.62 | |
| Video-ExpertTraining Strategy=Video-Expert2026.04 | 45.35 | |
| Image-ExpertTraining Strategy=Image-Expert2026.04 | 44.47 | |
| CoPDTraining Strategy=CoPD2026.04 | 43.77 | |
| MOPDTraining Strategy=MOPD2026.04 | 43.33 | |
| SFTDistillation Method=SFT, Teacher Model=Qwen3-32B2026.04 | 43.3 | |
| BaseTraining Strategy=Base2026.04 | 43.28 | |
| Video-Thinker-7BFrames=32, Protocol=think with video, Backbone=7B2026.05 | 43.1 | |
| HumanOmniV2Inference Mode=w/ Audio2026.04 | 42.9 | |
| Qwen3-VL-8B-Thinking + MVPFrames=1282026.01 | 42.6 | |
| Text-ExpertTraining Strategy=Text-Expert2026.04 | 42.19 | |
| Video-R1Inference Mode=w/o Audio2026.04 | 42.13 | |
| GPT-4oModel Category=Closed-source Large-sized MLLM2025.11 | 42 | |
| GPT-4oFrame Count=16-frame, Activation Replay=false2025.11 | 42 | |
| GPT-4oFrames=322026.05 | 42 | |
| Qwen3-VL-8B-ThinkingFrames=1282026.01 | 41.8 | |
| Claude-3.5Frame Count=16-frame, Activation Replay=false2025.11 | 41 | |
| Video-R1-7B + replayFrame Count=16-frame, Activation Replay=true2025.11 | 40.9 | |
| Omni-R1Inference Mode=w/ Audio2026.04 | 40.72 | |
| Qwen3-VL-8B-Thinking + MVPFrames=642026.01 | 40.1 | |
| InternVL3.5-8B + MVPFrames=642026.01 | 39.4 | |
| InternVL3.5-8BFrames=642026.01 | 39.3 | |
| HumanOmniV2Inference Mode=w/o Audio2026.04 | 38.87 | |
| InternVL3.5-8B + MVPFrames=1282026.01 | 38.8 | |
| Omni-R1Inference Mode=w/o Audio2026.04 | 38.76 | |
| Qwen3-VL-8B-ThinkingFrames=642026.01 | 38.5 | |
| STORMFrames=32, Backbone=Qwen2.5-VL-7B-Instruct, Sampling Strategy=uniform sampling2026.05 | 37.8 | |
| InternVL3.5-8BFrames=1282026.01 | 37.7 | |
| Jigsaw-7BFrames=1282026.01 | 37.6 | |
| Jigsaw-7BFrames=642026.01 | 36.9 | |
| Qwen2.5-VL-7B-Instruct + MVPFrames=1282026.01 | 36.7 | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 32026.05 | 36.6 | |
| VideoChat-R1-7BModel Category=Open-source Medium-sized MLLM2025.11 | 36.5 | |
| Video-R1-7BFrame Count=16-frame, Activation Replay=false2025.11 | 36.5 | |
| Video-R1-7BFrames=32, Backbone=7B, Sampling Strategy=uniform sampling2026.05 | 36.5 | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Frozen Questioner2026.05 | 36.3 | |
| Base ModelModel=Qwen3-VL-8B-Instruct, Training Iteration=w/o training2026.05 | 36.2 | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 12026.05 | 36 | |
| Qwen2.5-VL-7B-Instruct + MVPFrames=642026.01 | 35.9 | |
| LongVT-RFTFrames=32, Protocol=think with tools2026.05 | 35.5 | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 22026.05 | 35.2 | |
| Qwen2.5-VL-7B-InstructFrames=1282026.01 | 35 | |
| Video-R1-SFTDistillation Method=SFT, Teacher Model=Baseline2026.04 | 34.6 | |
| Qwen2.5-VL-7B-InstructFrames=642026.01 | 33.4 | |
| VideoChat-R1-7B (Alternative)Model Category=Open-source Medium-sized MLLM, Note=VideoChat-R1-7B [40]2025.11 | 33 | |
| InternVL-3-8BModel Category=Open-source Medium-sized MLLM2025.11 | 32.3 | |
| Qwen2.5-VL-7B-SFTFrames=32, Backbone=7B, Sampling Strategy=uniform sampling2026.05 | 31.4 | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 32026.05 | 31.1 | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Frozen Questioner2026.05 | 30.7 | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 12026.05 | 30.4 | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 22026.05 | 29.9 | |
| Base ModelModel=Qwen3-VL-4B-Instruct, Training Iteration=w/o training2026.05 | 29.9 | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 32026.05 | 29.3 | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 22026.05 | 29.1 | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 12026.05 | 28.6 | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Frozen Questioner2026.05 | 28.4 | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 12026.05 | 28 | |
| Qwen2-VL-7BModel Category=Open-source Medium-sized MLLM2025.11 | 27.8 | |
| Qwen2.5-VL-7BFrame Count=16-frame, Activation Replay=false2025.11 | 27.8 | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Frozen Questioner2026.05 | 27.8 | |
| Base ModelModel=Qwen2.5-VL-7B-Instruct, Training Iteration=w/o training2026.05 | 27.8 | |
| Qwen2.5-VL-7B-Instruct(CoT)Frames=32, Backbone=7B, Sampling Strategy=uniform sampling2026.05 | 27.8 | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 32026.05 | 27.2 | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 22026.05 | 26.9 | |
| Base ModelModel=Qwen2.5-VL-3B-Instruct, Training Iteration=w/o training2026.05 | 26.8 | |
| InternVL-2.5-8BModel Category=Open-source Medium-sized MLLM2025.11 | 23.6 |