Multi-discipline Reasoning on MMMU-Pro
52.2AccuracyLlama 4 Scout
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama 4 ScoutZero-shot=true2026.02 | 52.2 | |
| Qwen2.5-VL-32B + AT-RL (Ours)Zero-shot=true2026.02 | 51.9 | |
| OpenAI GPT-4oZero-shot=true2026.02 | 51.9 | |
| Gemini 2.0 FlashZero-shot=true2026.02 | 51.7 | |
| Qwen2.5-VL-72B InstructZero-shot=true2026.02 | 51.6 | |
| Claude 3.5 SonnetZero-shot=true2026.02 | 51.5 | |
| Qwen2.5-VL-32B + VPPOZero-shot=true2026.02 | 49.2 | |
| Qwen2.5-VL-32B InstructZero-shot=true2026.02 | 48.5 | |
| Qwen2.5-VL-7B-InstructModel scale=7B-8B2025.10 | 37.1 | |
| InternVL-3-8B-InstructModel scale=7B-8B2025.10 | 35.8 | |
| PLM-HoneyBee-8BModel scale=7B-8B2025.10 | 33.8 | |
| InternVL-2.5-8BModel scale=7B-8B2025.10 | 32 | |
| InternVL-2.5-4BModel scale=3B-4B2025.10 | 31.1 | |
| Qwen2.5-VL-3B-InstructModel scale=3B-4B2025.10 | 29.8 | |
| PLM-HoneyBee-3BModel scale=3B-4B2025.10 | 28.4 | |
| PLM-8BModel scale=7B-8B2025.10 | 20.5 | |
| PLM-3BModel scale=3B-4B2025.10 | 19.5 | |
| PLM-HoneyBee-1BModel scale=1B2025.10 | 18.8 | |
| InternVL-2.5-1BModel scale=1B2025.10 | 16.2 | |
| InternVL-3-1B-InstructModel scale=1B2025.10 | 16.2 | |
| PLM-1BModel scale=1B2025.10 | 15.8 |