Text-centric Reasoning on VideoThinkBench mini (test)
89Average ScoreGemini 2.5 Pro
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 2.5 ProModel Category=Vision-Language Models2025.11 | 89 | 100 | 100 | 100 | 90 | 100 | 83.3 | 93.3 | 80 | 80 | 85 | 65 | 95 | 85 | |
| GPT-5 highModel Category=Vision-Language Models2025.11 | 86.6 | 100 | 100 | 100 | 100 | 100 | 83.3 | 93.3 | 80 | 83.9 | 75 | 55 | 85 | 70 | |
| Qwen3-VL-235B-A22BModel Category=Vision-Language Models2025.11 | 77.6 | 100 | 100 | 80 | 50 | 84.6 | 58.3 | 100 | 56 | 80 | 70 | 65 | 90 | 75 | |
| Claude Sonnet 4.5Model Category=Vision-Language Models2025.11 | 77.2 | 100 | 100 | 60 | 40 | 100 | 83.3 | 100 | 60 | 80 | 80 | 45 | 80 | 75 | |
| Qwen3-VL-PlusModel Category=Vision-Language Models2025.11 | 75.8 | 100 | 95 | 100 | 70 | 76.9 | 66.7 | 80 | 64 | 57.1 | 65 | 65 | 80 | 65 | |
| Qwen3-VL-32BModel Category=Vision-Language Models2025.11 | 72.5 | 100 | 95 | 80 | 50 | 76.9 | 66.7 | 93.3 | 40 | 65.7 | 75 | 45 | 90 | 65 | |
| Sora-2Model Category=Video Generation Models, Input Modality=Audio2025.11 | 67.6 | 100 | 90 | 50 | 40 | 76.9 | 66.7 | 73.3 | 56 | 45.7 | 75 | 45 | 90 | 70 | |
| Nano Banana ProModel Category=Image Generation Models2025.11 | 66 | 56.7 | 65 | 80 | 80 | 69.2 | 75 | 80 | 44 | 62.9 | 75 | 45 | 75 | 50 | |
| Sora-2Model Category=Video Generation Models, Input Modality=Last Frame2025.11 | 57.1 | 76.7 | 65 | 40 | 30 | 69.2 | 66.7 | 73.3 | 52 | 54.3 | 70 | 45 | 60 | 40 | |
| Seedream 4.5Model Category=Image Generation Models2025.11 | 55.7 | 100 | 80 | 20 | 10 | 69.2 | 75 | 60 | 36 | 48.6 | 55 | 60 | 55 | 55 | |
| Veo 3.1Model Category=Video Generation Models, Input Modality=Last Frame2025.11 | 48.3 | 80 | 70 | 50 | 20 | 61.5 | 16.7 | 60 | 52 | 42.9 | 50 | 45 | 35 | 45 | |
| Veo 3.1Model Category=Video Generation Models, Input Modality=Audio2025.11 | 44.5 | 93.3 | 80 | 50 | 20 | 61.5 | 41.7 | 80 | 40 | 51.4 | 25 | 5 | 20 | 10 | |
| GPT Image 1.5Model Category=Image Generation Models2025.11 | 41.4 | 90 | 40 | 0 | 0 | 69.2 | 25 | 46.7 | 40 | 22.9 | 50 | 40 | 65 | 50 | |
| MiniMax Hailuo 2.3Model Category=Video Generation Models2025.11 | 38.4 | 76.6 | 40 | 10 | 20 | 61.5 | 33.3 | 86.6 | 16 | 65.7 | 30 | 10 | 30 | 20 | |
| MOVA-720pModel Category=Video Generation Models, Input Modality=Last Frame2025.11 | 12.5 | 30 | 35 | 20 | 0 | 15.4 | 0 | 6.7 | 8 | 2.9 | 0 | 25 | 10 | 10 | |
| Qwen-Image-Edit-2511Model Category=Image Generation Models2025.11 | 10.5 | 0 | 0 | 0 | 0 | 15.4 | 16.7 | 6.7 | 4 | 8.6 | 15 | 15 | 30 | 25 | |
| MOVA-360pModel Category=Video Generation Models, Input Modality=Last Frame2025.11 | 10.4 | 20 | 10 | 0 | 20 | 15.4 | 0 | 0 | 20 | 0 | 5 | 30 | 5 | 10 | |
| BAGELModel Category=Image Generation Models, Output Modality=Image Output2025.11 | 8.9 | 6.7 | 0 | 0 | 0 | 0 | 0 | 6.7 | 4 | 2.9 | 25 | 40 | 10 | 20 | |
| MOVA-720pModel Category=Video Generation Models, Input Modality=Audio2025.11 | 7.2 | 0 | 0 | 0 | 0 | 30.8 | 16.7 | 0 | 8 | 2.9 | 0 | 10 | 20 | 5 | |
| MOVA-360pModel Category=Video Generation Models, Input Modality=Audio2025.11 | 5.8 | 0 | 0 | 0 | 0 | 30.8 | 33.3 | 0 | 0 | 5.7 | 0 | 0 | 5 | 0 | |
| Seedance 1.0 ProModel Category=Video Generation Models2025.11 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | |
| Wan2.2-TI2V-5BModel Category=Video Generation Models2025.11 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |