Multimodal Reasoning on MMMUPro (Pass@1)
43.47Pass@1Self-distill Masking-KD-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| Self-distill Masking-KD-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=4096, Distillation=Self-distill2026.05 | 43.47 | |
| Masking-KD-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=4096, Distillation=Distilled from 8B2026.05 | 40.52 | |
| Qwen3-VL-8B-ThinkingModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 39.83 | |
| InternVL3.5-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 38.5 | |
| InternVL3-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 36.01 | |
| MiMo-VL-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 34.57 | |
| Qwen3-VL-4B-ThinkingModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 32.08 | |
| Qwen2.5-VL-7BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 31.56 | |
| Masking-KD-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=4096, Distillation=Distilled from 8B2026.05 | 30.75 | |
| InternVL3.5-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 29.77 | |
| Qwen2.5-VL-3BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 27.75 | |
| InternVL3-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 22.95 | |
| Ovis2-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 16.18 | |
| Qwen3-VL-2B-ThinkingModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 14.51 | |
| Ovis2-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 12.6 | |
| InternVL3.5-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 12.54 | |
| Ovis2-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 10.23 |