General Visual Reasoning on MMMU
69.1AccuracyGPT-4o
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4o2026.02 | 69.1 | — | |
| Claude-3.5-Sonnet2026.02 | 68.3 | — | |
| Supervised ETTCEnsemble Composition=Same Family Models2026.05 | 66.46 | 1.12 | |
| ETTCEnsemble Composition=Same Family Models2026.05 | 65.34 | — | |
| Supervised ETTCEnsemble Composition=Similar Size Models2026.05 | 59.01 | 0.38 | |
| ETTCEnsemble Composition=Similar Size Models2026.05 | 58.63 | — | |
| VotingEnsemble Composition=Same Family Models2026.05 | 58.63 | — | |
| Qwen2.5-VL-32BParameters=32B2026.02 | 57.44 | — | |
| InternVL2.5-38BParameters=38B2026.02 | 56.98 | — | |
| RuCL2026.02 | 56.67 | — | |
| ThinkLite-VL-7BParameters=7B2026.02 | 55.44 | — | |
| VL-Rethinker-7BParameters=7B2026.02 | 54.67 | — | |
| OpenVLThinker-7BParameters=7B2026.02 | 54.29 | — | |
| MM-Eureka-7BParameters=7B2026.02 | 53.78 | — | |
| VotingEnsemble Composition=Similar Size Models2026.05 | 53.66 | — | |
| Perception-R1-7BParameters=7B2026.02 | 53.11 | — | |
| Average (Single-Model)Ensemble Composition=Same Family Models2026.05 | 52.79 | — | |
| Qwen2.5-VL-7BParameters=7B2026.02 | 51 | — | |
| Average (Single-Model)Ensemble Composition=Similar Size Models2026.05 | 48.39 | — | |
| InternVL2.5-8BParameters=8B2026.02 | 45.73 | — | |
| R1-Onevision-7BParameters=7B2026.02 | 43.7 | — | |
| Vision-R1-7BParameters=7B2026.02 | 43.28 | — |