Multi-modal Reasoning on MMMU-Pro
85.6AccuracyCoT2-Meta
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CoT2-MetaBackbone=Claude-4.5, Strategy=Ours (CoT2-Meta)2026.03 | 85.6 | 14.5 | |
| CoT2-MetaBackbone=DeepSeek-V3.2, Strategy=Ours (CoT2-Meta)2026.03 | 78.4 | 10.5 | |
| Vanilla ToTBackbone=Claude-4.5, Strategy=Vanilla ToT2026.03 | 77.8 | 8.3 | |
| Best-of-16Backbone=Claude-4.5, Strategy=Best-of-162026.03 | 73.1 | 4.8 | |
| Vanilla ToTBackbone=DeepSeek-V3.2, Strategy=Vanilla ToT2026.03 | 71.5 | 5.7 | |
| Greedy CoTBackbone=Claude-4.5, Strategy=Greedy CoT2026.03 | 68.4 | — | |
| Best-of-16Backbone=DeepSeek-V3.2, Strategy=Best-of-162026.03 | 68.2 | 3 | |
| Greedy CoTBackbone=DeepSeek-V3.2, Strategy=Greedy CoT2026.03 | 64.6 | — | |
| AutoNPO2026.04 | 57.24 | — | |
| NPOstage=early + late-stage2026.04 | 57.07 | — | |
| NPOstage=early-stage only2026.04 | 56.85 | — | |
| GRPOtype=pure on-policy2026.04 | 55.78 | — | |
| ExGRPOtype=historical replay2026.04 | 55.49 | — | |
| RLEPtype=far future2026.04 | 55.38 | — | |
| CoT2-MetaBackbone=Qwen2.5-VL-7B, Strategy=Ours (CoT2-Meta)2026.03 | 55.2 | 12.2 | |
| LUFFYtype=external teacher2026.04 | 54.23 | — | |
| Qwen3-VL-8B-Instruct2026.04 | 51.75 | — | |
| Vanilla ToTBackbone=Qwen2.5-VL-7B, Strategy=Vanilla ToT2026.03 | 48.6 | 6.4 | |
| Best-of-16Backbone=Qwen2.5-VL-7B, Strategy=Best-of-162026.03 | 44.5 | 3.4 | |
| Greedy CoTBackbone=Qwen2.5-VL-7B, Strategy=Greedy CoT2026.03 | 41 | — | |
| Qwen2.5-VL-7B + PGPOModel Size=7B, Optimization Strategy=PGPO2026.04 | 39.01 | — | |
| Qwen2.5-VL-7B + VPPOModel Size=7B, Optimization Strategy=VPPO2026.04 | 38.75 | — | |
| Qwen2.5-VL-7B + PAPOModel Size=7B, Optimization Strategy=PAPO2026.04 | 38.38 | — | |
| VL-Rethinker-7BModel Size=7B2026.04 | 37.13 | — | |
| Qwen2.5-VL-7B + DAPOModel Size=7B, Optimization Strategy=DAPO2026.04 | 36.99 | — | |
| Qwen2.5-VL-7B + GRPOModel Size=7B, Optimization Strategy=GRPO2026.04 | 36.84 | — | |
| R1-ShareVL-7BModel Size=7B2026.04 | 35.1 | — | |
| NoisyRollout-7BModel Size=7B2026.04 | 34.5 | — | |
| MM-Eureka-7BModel Size=7B2026.04 | 30.3 | — | |
| Qwen2.5-VL-3B + PGPOModel Size=3B, Optimization Strategy=PGPO2026.04 | 29.33 | — | |
| Qwen2.5-VL-3B + VPPOModel Size=3B, Optimization Strategy=VPPO2026.04 | 28.83 | — | |
| Qwen2.5-VL-3B + PAPOModel Size=3B, Optimization Strategy=PAPO2026.04 | 28.76 | — | |
| Qwen2.5-VL-3B + GRPOModel Size=3B, Optimization Strategy=GRPO2026.04 | 28.02 | — | |
| Qwen2.5-VL-3B + DAPOModel Size=3B, Optimization Strategy=DAPO2026.04 | 27.2 | — | |
| Qwen2.5-VL-7BModel Size=7B2026.04 | 26.67 | — | |
| Qwen2.5-VL-3BModel Size=3B2026.04 | 20.95 | — |