Multi-modal Evaluation on MME-RW
31.9Mean AccuracyTTAug
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TTAugAdaptation strategy=Test-time Augmentation2025.10 | 31.9 | — | — | — | |
| (2)Adaptation strategy=Model parameter adaptation2025.10 | 31.4 | — | — | — | |
| TTAugtest-time scaling=Method 52025.10 | 31.1 | — | — | — | |
| (1)Adaptation strategy=(1)2025.10 | 30.9 | — | — | — | |
| Baselinetest-time scaling=none2025.10 | 27.8 | — | — | — | |
| Baseline2025.10 | 27.8 | — | — | — | |
| Method ③test-time scaling=Other method 32025.10 | 27.6 | — | — | — | |
| Method ④test-time scaling=Other method 42025.10 | 27.6 | — | — | — | |
| Method ②test-time scaling=Other method 22025.10 | 26.4 | — | — | — | |
| Method ①test-time scaling=Other method 12025.10 | 26.2 | — | — | — | |
| Adaptive-CoFReasoning Paradigm=Thinking with Images, Base Model=Qwen2.5-VL-7B2026.02 | — | 50.9 | — | — | |
| DeepEyesReasoning Paradigm=Thinking with Images, Base Model=Qwen2.5-VL-7B2026.02 | — | 49.5 | — | — | |
| GPT-4oModel=GPT-4o2025.09 | — | 45.2 | 46.4 | 42.3 | |
| HiDeModel=Qwen2.5-VL 7B2025.09 | — | 63.8 | 66.7 | 42.9 | |
| MIRRORReasoning Paradigm=Thinking with Images, Base Model=Qwen2.5-VL-7B2026.02 | — | 51.49 | — | — | |
| PixelReasonerReasoning Paradigm=Thinking with Images, Base Model=Qwen2.5-VL-7B2026.02 | — | 49.7 | — | — | |
| Qwen2.5-VL 7BModel=Qwen2.5-VL 7B2025.09 | — | 61.4 | 64.3 | 40.1 | |
| VicropModel=Qwen2.5-VL 7B2025.09 | — | 62.3 | 65.1 | 42 | |
| VL-RethinkerReasoning Paradigm=Text Reflection, Base Model=Qwen2.5-VL-7B2026.02 | — | 47.21 | — | — |