Multi-discipline Reasoning on EMMA core
24.6AccuracyLlama 4 Scout
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama 4 ScoutZero-shot=true2026.02 | 24.6 | |
| Qwen2.5-VL-32B + AT-RL (Ours)Zero-shot=true2026.02 | 19.4 | |
| Claude 3.5 SonnetZero-shot=true2026.02 | 18.7 | |
| Qwen2.5-VL-32B + VPPOZero-shot=true2026.02 | 17.8 | |
| Qwen2.5-VL-72B InstructZero-shot=true2026.02 | 17.7 | |
| Gemini 2.0 FlashZero-shot=true2026.02 | 17.2 | |
| Qwen2.5-VL-32B InstructZero-shot=true2026.02 | 14.7 | |
| OpenAI GPT-4oZero-shot=true2026.02 | 6.3 |