Multimodal Reasoning on MMMU-Pro (Acc, # Tokens)
58.4AccuracyLaRe
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LaReBase VLM=Qwen3-VL-4B-Instruct2025.11 | 58.4 | 103.5 | |
| CCoTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 55.5 | 234.9 | |
| LVR*Base VLM=Qwen3-VL-4B-Instruct2025.11 | 53.9 | 144.4 | |
| CoF*Base VLM=Qwen3-VL-4B-Instruct2025.11 | 53.7 | 269.1 | |
| Standard CoTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 52.4 | 226.8 | |
| ICoTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 52.3 | 211.7 | |
| MCOUTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 52 | 124.5 | |
| MM-CoTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 49.2 | 204.2 | |
| LaReBase VLM=Qwen3-VL-2B-Instruct2025.11 | 44.2 | 102.4 | |
| STRIDEBase model=Qwen2.5-VL-7B2026.06 | 44.2 | — | |
| SRPOBase model=Qwen2.5-VL-7B2026.06 | 42.3 | — | |
| CoF*Base VLM=Qwen3-VL-2B-Instruct2025.11 | 41.2 | 274 | |
| ICoTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 40.8 | 214.7 | |
| MM-CoTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 39.8 | 198.8 | |
| MCOUTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 39 | 120.1 | |
| LVR*Base VLM=Qwen3-VL-2B-Instruct2025.11 | 38.9 | 134.7 | |
| MM-EurekaBase model=Qwen2.5-VL-7B2026.06 | 37.6 | — | |
| Standard CoTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 37 | 224.2 | |
| CCoTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 35.5 | 203.5 | |
| RECAPBackbone=Qwen2.5-VL-7B, Training=finetuned2025.10 | 34.15 | — | |
| Reasoning-onlyBackbone=Qwen2.5-VL-7B, Training=finetuned2025.10 | 33.87 | — | |
| UniformBackbone=Qwen2.5-VL-7B, Training=finetuned2025.10 | 31.91 | — | |
| CoresetBackbone=Qwen2.5-VL-7B, Training=finetuned2025.10 | 31.91 | — | |
| PropMixBackbone=Qwen2.5-VL-7B, Training=finetuned2025.10 | 31.39 | — | |
| BaseBase model=Qwen2.5-VL-7B2026.06 | 30.5 | — | |
| LwFBackbone=Qwen2.5-VL-7B, Training=finetuned2025.10 | 29.59 | — | |
| CFPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 27.96 | — | |
| Vision-R1-7B2025.10 | 26.76 | — | |
| VLAA-Thinker-7B2025.10 | 26.3 | — | |
| Qwen2.5-VL-7Bvariant=Base model2025.10 | 25.55 | — | |
| PAPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 23.42 | — | |
| OpenVLThinker-7B2025.10 | 21.79 | — | |
| GRPOBackbone=Qwen3-VL-2B-Thinking2026.06 | 20.51 | — |