Multimodal Reasoning on MMK12 (Pass@1)
51Pass@1Masking-KD-4B
Evaluation Results
| Method | Links | |
|---|---|---|
| Masking-KD-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=4096, Distillation=Distilled from 8B2026.05 | 51 | |
| Self-distill Masking-KD-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=4096, Distillation=Self-distill2026.05 | 49.95 | |
| Ovis2-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 48.15 | |
| InternVL3.5-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 44.95 | |
| MiMo-VL-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 44.1 | |
| Qwen2.5-VL-7BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 43.9 | |
| Qwen3-VL-8B-ThinkingModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 42.55 | |
| Qwen2.5-VL-3BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 39.85 | |
| InternVL3-8BModel Size=~8B, Decoding Strategy=greedy, Max Length=40962026.05 | 39.8 | |
| Ovis2-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 39.1 | |
| Masking-KD-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=4096, Distillation=Distilled from 8B2026.05 | 37.2 | |
| InternVL3-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 37 | |
| Ovis2-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 32.45 | |
| Qwen3-VL-4B-ThinkingModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 31.55 | |
| InternVL3.5-4BModel Size=~4B, Decoding Strategy=greedy, Max Length=40962026.05 | 25.8 | |
| InternVL3.5-2BModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 14.7 | |
| Qwen3-VL-2B-ThinkingModel Size=~2B, Decoding Strategy=greedy, Max Length=40962026.05 | 13 |