Mathematical multi-modal reasoning on WeMath
85.11Pass@1Qwen3-VL-8B-DeepVision
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL-8B-DeepVisionSeries=Qwen3-VL-8B, Training Algorithm=GSPO2026.02 | 85.11 | |
| Qwen3-VL-8B-ThinkingSeries=Qwen3-VL-8B2026.02 | 84.54 | |
| Gemini-2.5-Flash-LiteModel Category=Closed-source Models2026.02 | 83.85 | |
| MiMo-VL-7B-OpenMMReasonerSeries=MiMo-VL-7B2026.02 | 83.45 | |
| MiMo-VL-7B-DeepVisionSeries=MiMo-VL-7B, Training Algorithm=GSPO2026.02 | 82.98 | |
| Qwen3-VL-8B-InstructSeries=Qwen3-VL-8B2026.02 | 79.36 | |
| MiMo-VL-7B-MM-EurekaSeries=MiMo-VL-7B2026.02 | 79.08 | |
| GPT-5-Nano-HighModel Category=Closed-source Models2026.02 | 78.62 | |
| MiMo-VL-7B-MathBookSeries=MiMo-VL-7B2026.02 | 77.18 | |
| MiMo-VL-7B-RL-2508Series=MiMo-VL-7B2026.02 | 76.95 | |
| MiMo-VL-7B-SFT-2508Series=MiMo-VL-7B2026.02 | 74.42 | |
| Self-distill Masking-KD-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=4096, Distillation=Self-distill2026.05 | 71.72 | |
| Masking-KD-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=4096, Distillation=Distilled from 8B2026.05 | 71.03 | |
| Qwen3-VL-8B-ThinkingModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 66.15 | |
| Ovis2-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 64.66 | |
| Masking-KD-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=4096, Distillation=Distilled from 8B2026.05 | 63.79 | |
| Ovis2-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 60.29 | |
| InternVL3.5-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 56.61 | |
| Ovis2-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 51.95 | |
| InternVL3-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 51.32 | |
| Qwen2.5-VL-3BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 49.66 | |
| Qwen3-VL-4B-ThinkingModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 49.37 | |
| Qwen2.5-VL-7BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 48.74 | |
| InternVL3-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 47.93 | |
| InternVL3.5-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 45.46 | |
| MiMo-VL-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 42.41 | |
| Qwen2.5-VL-7b-RAPBackbone=Qwen2.5-VL-7b, Sample=5,159, Time (h)=52.82025.06 | 42 | |
| Qwen2.5-VL-3b-RAPBackbone=Qwen2.5-VL-3b, Sample=4,374, Time (h)=322025.06 | 29.33 | |
| Qwen3-VL-2B-ThinkingModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 25.17 | |
| InternVL3.5-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 24.31 |