Multimodal Reasoning on M3CoT (test)
91.61Total AccHuman
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HumanProtocol=Human2025.12 | 91.61 | 97.83 | 92.62 | 94.31 | 96.28 | 92.41 | 88.71 | 87.23 | 88.75 | 85.71 | — | — | |
| OursModel=Qwen2-VL 7B2026.04 | 73 | — | — | — | — | — | — | — | — | — | 7 | 0.86 | |
| IVT-LRModel=Qwen2-VL 7B2026.04 | 69.8 | — | — | — | — | — | — | — | — | — | 10 | 0.67 | |
| Chain-of-FocusModel=Qwen2-VL 7B2026.04 | 64.3 | — | — | — | — | — | — | — | — | — | 185.7 | 2.63 | |
| Ours [with DPS]Model Category=Reasoning MLLMs, Training Strategy=DPS2026.01 | 63.9 | — | — | — | — | — | — | — | — | — | — | — | |
| OursModel Category=Reasoning MLLMs, Training Strategy=DAPO2026.01 | 63.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Ours [with DPS and annealing]Model Category=Reasoning MLLMs, Training Strategy=Two-stage RL2026.01 | 62.7 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4VModel Category=Closed-Source MLLMs2026.01 | 62.6 | — | — | — | — | — | — | — | — | — | — | — | |
| MIND_largeModel Size=738M, Protocol=Finetuning2025.12 | 61.56 | 79.62 | 66.41 | 48.57 | 81.11 | 69.42 | 51.22 | 50.71 | 61.25 | 47.62 | — | — | |
| VLAA-Thinker 7BModel Category=Reasoning MLLMs2026.01 | 61.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5VL 7BModel Category=Open-Source General MLLMs2026.01 | 60.1 | — | — | — | — | — | — | — | — | — | — | — | |
| MixedR1 7BModel Category=Reasoning MLLMs2026.01 | 59.9 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-V1.5Model Size=13B, Protocol=Finetuning2025.12 | 59.5 | 68.72 | 72.41 | 40.86 | 83.52 | 64.61 | 69.11 | 35.71 | 45 | 38.1 | — | — | |
| CogVLMModel Size=17B, Protocol=Finetuning2025.12 | 58.25 | 65.88 | 77.52 | 29.09 | 81.32 | 65.43 | 75.61 | 35.71 | 46.25 | 47.62 | — | — | |
| MC-CoT_largeModel Size=738M, Protocol=Finetuning2025.12 | 57.69 | 42.65 | 67.43 | 50.56 | 58.24 | 60.49 | 56.1 | 57.86 | 62.5 | 14.29 | — | — | |
| MIND_baseModel Size=223M, Protocol=Finetuning2025.12 | 57.38 | 72.04 | 63.09 | 43.31 | 82.22 | 65.29 | 44.72 | 49.29 | 61.25 | 33.33 | — | — | |
| R1-Onevision 7BModel Category=Reasoning MLLMs2026.01 | 57.3 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT4VProtocol=Zero-shot2025.12 | 56.95 | 80.09 | 54.66 | 43.95 | 87.78 | 67.77 | 82.11 | 42.14 | 43.75 | 42.86 | — | — | |
| GPT-4oModel Category=Closed-Source MLLMs2026.01 | 55.7 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaMA-AdaperModel Size=7B, Protocol=Finetuning2025.12 | 54.89 | 62.56 | 72.29 | 30.21 | 76.92 | 59.67 | 72.36 | 30.71 | 38.75 | 38.1 | — | — | |
| MC-CoT_baseModel Size=223M, Protocol=Finetuning2025.12 | 53.51 | 53.55 | 63.98 | 43.56 | 61.54 | 69.55 | 29.27 | 42.86 | 33.75 | 28.57 | — | — | |
| VisonR1 7BModel Category=Reasoning MLLMs2026.01 | 53.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Multimodal-CoT_largeModel Size=738M, Protocol=Finetuning2025.12 | 48.73 | 45.5 | 50.19 | 43.56 | 63.74 | 64.61 | 33.33 | 40.71 | 61.25 | 28.57 | — | — | |
| ICoTModel=Qwen2-VL 7B2026.04 | 46 | — | — | — | — | — | — | — | — | — | 96.5 | 2.86 | |
| CSMRPerception Backbone=Qwen2-VL-7B, Reasoning Backbone=Qwen2-7B, Zero-shot=true2026.05 | 45.7 | — | — | — | — | — | — | — | — | — | — | — | |
| No-CoTModel=Qwen2-VL 7B2026.04 | 45.4 | — | — | — | — | — | — | — | — | — | — | — | |
| GeminiProtocol=Zero-shot2025.12 | 45.17 | 73.93 | 41.25 | 31.21 | 56.67 | 71.49 | 62.6 | 30.71 | 27.5 | 28.57 | — | — | |
| SCAFFOLDModel=Qwen2-VL 7B2026.04 | 44.9 | — | — | — | — | — | — | — | — | — | 170.8 | 5.14 | |
| Multimodal-CoT_baseModel Size=223M, Protocol=Finetuning2025.12 | 44.85 | 41.71 | 46.49 | 39.9 | 59.34 | 60.91 | 27.64 | 48.57 | 35 | 28.57 | — | — | |
| CCoTModel=Qwen2-VL 7B2026.04 | 44.1 | — | — | — | — | — | — | — | — | — | 177.2 | 5.31 | |
| ICoTPerception Backbone=Qwen2-VL-7B, Reasoning Backbone=Qwen2-VL-7B, Zero-shot=true2026.05 | 44.1 | — | — | — | — | — | — | — | — | — | — | — | |
| No-CoTPerception Backbone=Qwen2-VL-7B, Reasoning Backbone=Qwen2-VL-7B, Zero-shot=true2026.05 | 43.6 | — | — | — | — | — | — | — | — | — | — | — | |
| OursModel=Chameleon 7B2026.04 | 43.4 | — | — | — | — | — | — | — | — | — | 7 | 1.24 | |
| CCoTPerception Backbone=Qwen2-VL-7B, Reasoning Backbone=Qwen2-VL-7B, Zero-shot=true2026.05 | 43.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Multimodal CoTModel=Qwen2-VL 7B2026.04 | 42.5 | — | — | — | — | — | — | — | — | — | 106.3 | 3.1 | |
| SCAFFOLDPerception Backbone=Qwen2-VL-7B, Reasoning Backbone=Qwen2-VL-7B, Zero-shot=true2026.05 | 41.7 | — | — | — | — | — | — | — | — | — | — | — | |
| CaptionPerception Backbone=Qwen2-VL-7B, Reasoning Backbone=Qwen2-7B, Zero-shot=true2026.05 | 40.9 | — | — | — | — | — | — | — | — | — | — | — | |
| IVT-LRModel=Chameleon 7B2026.04 | 40.8 | — | — | — | — | — | — | — | — | — | 10 | 1.13 | |
| Multimodal CoTPerception Backbone=Qwen2-VL-7B, Reasoning Backbone=Qwen2-VL-7B, Zero-shot=true2026.05 | 40.1 | — | — | — | — | — | — | — | — | — | — | — | |
| DDCoTPerception Backbone=Qwen2-VL-7B, Reasoning Backbone=Qwen2-7B, Zero-shot=true2026.05 | 39 | — | — | — | — | — | — | — | — | — | — | — | |
| CogVLMModel Size=17B, Protocol=Zero-shot2025.12 | 37.19 | 52.61 | 37.42 | 26.91 | 55.56 | 54.13 | 29.27 | 29.29 | 32.5 | 23.81 | — | — | |
| Chain-of-FocusModel=Chameleon 7B2026.04 | 36.5 | — | — | — | — | — | — | — | — | — | 739.4 | 3.09 | |
| InstructBLIPModel Size=13B, Protocol=Zero-shot2025.12 | 35.94 | 38.39 | 30.52 | 26.27 | 76.67 | 70.66 | 35.77 | 30 | 22.5 | 19.05 | — | — | |
| ChameleonProtocol=Tool-Usage2025.12 | 34.29 | 43.87 | 26.05 | 25.44 | 39.13 | 37.3 | 48.78 | 17.73 | 26.25 | 23.81 | — | — | |
| ICoTModel=Chameleon 7B2026.04 | 32.3 | — | — | — | — | — | — | — | — | — | 110.9 | 5.43 | |
| IdealGPTProtocol=Tool-Usage2025.12 | 32.19 | 31.73 | 31.63 | 26.23 | 56.52 | 50 | 26.83 | 20.57 | 30 | 38.1 | — | — | |
| CCoTModel=Chameleon 7B2026.04 | 31.4 | — | — | — | — | — | — | — | — | — | 168.4 | 5.35 | |
| SCAFFOLDModel=Chameleon 7B2026.04 | 31.1 | — | — | — | — | — | — | — | — | — | 194.3 | 6.12 | |
| Multimodal CoTModel=Chameleon 7B2026.04 | 30.6 | — | — | — | — | — | — | — | — | — | 110.5 | 3.62 | |
| RandomProtocol=Random2025.12 | 28.56 | 32.7 | 30.62 | 26.71 | 32.97 | 22.22 | 20.33 | 35.71 | 27.5 | 23.81 | — | — | |
| No-CoTModel=Chameleon 7B2026.04 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | |
| LLava-V1.5Model Size=13B, Protocol=Zero-shot2025.12 | 27.05 | 36.97 | 27.46 | 20.22 | 52.22 | 23.55 | 27.64 | 22.86 | 45 | 4.76 | — | — | |
| VisualChatGPTModel Size=>175B, Protocol=Tool-Usage2025.12 | 25.92 | 30.09 | 36.28 | 7.78 | 43.48 | 29.92 | 33.33 | 21.99 | 21.25 | 28.57 | — | — | |
| Kosmos-2Model Size=2B, Protocol=Zero-shot2025.12 | 23.17 | 10.43 | 28.61 | 21.18 | 33.33 | 17.77 | 28.46 | 21.43 | 21.25 | 14.29 | — | — | |
| HuggingGPTModel Size=175B, Protocol=Tool-Usage2025.12 | 14.6 | 17.57 | 20.93 | 10.33 | 8.7 | 14.75 | 9.76 | 11.35 | 22.5 | 9.52 | — | — |