Multi-modal Reasoning on MMK12 (test)
73.9Accuracyo1
Evaluation Results
| Method | Links | |
|---|---|---|
| o12026.03 | 73.9 | |
| Qwen-2.5-VL-72B2026.03 | 70.5 | |
| DIVA-GRPO-7B2026.03 | 70.2 | |
| R1-ShareVL-7B2026.03 | 68.8 | |
| Qwen-2.5-VL-32B2026.03 | 66.8 | |
| Gemini2-flash2026.03 | 65.2 | |
| MM-Eureka-7B2026.03 | 64.5 | |
| InternVL2.5-VL-78B2026.03 | 61.6 | |
| QVQ-72B-Preview2026.03 | 61.5 | |
| OpenVLThinker-7B2026.03 | 60.6 | |
| Adora-7B2026.03 | 58.1 | |
| InternVL2.5-VL-38B2026.03 | 58 | |
| Claude3.7-Sonnet2026.03 | 55.3 | |
| SFT-7B2026.03 | 54.3 | |
| Qwen-2.5-VL-7B2026.03 | 53.6 | |
| SFT-MModel=Qwen2.5-VL-7B, Paradigm=SFT-M, Decoding strategy=Greedy, Temperature=0, Repetition penalty=1.1, Max tokens=40962026.02 | 50.7 | |
| GPT-4o2026.03 | 49.9 | |
| SFTModel=Qwen2.5-VL-7B, Paradigm=SFT, Decoding strategy=Greedy, Temperature=0, Repetition penalty=1.1, Max tokens=40962026.02 | 49.2 | |
| GRPOModel=Qwen2.5-VL-7B, Paradigm=GRPO, Decoding strategy=Greedy, Temperature=0, Repetition penalty=1.1, Max tokens=40962026.02 | 49.05 | |
| SFT-RSModel=Qwen2.5-VL-7B, Paradigm=SFT-RS, Decoding strategy=Greedy, Temperature=0, Repetition penalty=1.1, Max tokens=40962026.02 | 48.6 | |
| InternVL2.5-38B-MPO2026.03 | 48.3 | |
| InternVL2.5-VL-8B2026.03 | 45.6 | |
| SFT-MModel=Qwen2.5-VL-3B, Paradigm=SFT-M, Decoding strategy=Greedy, Temperature=0, Repetition penalty=1.1, Max tokens=40962026.02 | 43.4 | |
| SFTModel=Qwen2.5-VL-3B, Paradigm=SFT, Decoding strategy=Greedy, Temperature=0, Repetition penalty=1.1, Max tokens=40962026.02 | 42.2 | |
| GRPOModel=Qwen2.5-VL-3B, Paradigm=GRPO, Decoding strategy=Greedy, Temperature=0, Repetition penalty=1.1, Max tokens=40962026.02 | 42.1 | |
| SFT-RSModel=Qwen2.5-VL-3B, Paradigm=SFT-RS, Decoding strategy=Greedy, Temperature=0, Repetition penalty=1.1, Max tokens=40962026.02 | 41.8 | |
| R1-Onevision-7B2026.03 | 39.8 | |
| InternVL2.5-8B-MPO2026.03 | 34.5 |