Real-world Multimodal Reasoning on RealWorldQA
75.4AccuracyGPT-4o
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4oZero-shot=true, Decoding Strategy=greedy2024.09 | 75.4 | — | |
| GPT-4o2024.09 | 75.4 | — | |
| PerceptionLM-8Bmask=12025.10 | 75 | — | |
| EMOVAModel Size=72B2024.09 | 71 | — | |
| MM1.5-30BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 69 | — | |
| Grok-1.5Vopen-source=false2024.04 | 68.7 | — | |
| Gemini Pro 1.52024.09 | 68.7 | — | |
| Cambrian-34BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 67.8 | — | |
| Gemini Pro 1.5open-source=false2024.04 | 67.5 | — | |
| InternVL 1.2#param=40B, open-source=false2024.04 | 67.5 | — | |
| EMOVAModel Size=7B2024.09 | 67.5 | — | |
| VITAModel Size=1.52024.09 | 66.8 | — | |
| InternVL 1.5#param=26B, open-source=false2024.04 | 66 | — | |
| SINKTRACKBase LLM=Qwen2.5-VL-7B-Instruct2026.04 | 65.49 | 52.81 | |
| Gemini-1.5-ProZero-shot=true, Decoding Strategy=greedy2024.09 | 64.1 | — | |
| EMOVAModel Size=3B2024.09 | 62.6 | — | |
| Baichuan-OmniModel Size=7B2024.09 | 62.6 | — | |
| MM1.5-7BModel Scale=7B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 62.5 | — | |
| GAR-8Bmask=1, Training Data=w/ 600K General Data2025.10 | 61.8 | — | |
| GPT-4Vopen-source=false2024.04 | 61.4 | — | |
| GPT-4V2024.09 | 61.4 | — | |
| MM1.5-3B-MoEModel Scale=3B, Architecture=MoE, Zero-shot=true, Decoding Strategy=greedy2024.09 | 60.7 | — | |
| BLIP-3Model Scale=3B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 60.5 | — | |
| H-GRPOBackbone=Qwen2.5-VL-3B2026.06 | 60.3 | — | |
| Phi-3-Vision-4BModel Scale=3B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 59.4 | — | |
| MM1-30BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 59.4 | — | |
| VITAModel Size=8x7B2024.09 | 59 | — | |
| Qwen2.5-VLParameters=3B2026.06 | 59 | — | |
| GAR-8Bmask=12025.10 | 58.7 | — | |
| SINKTRACKBase LLM=Gemma3-12B-Instruct2026.04 | 57.91 | 52.54 | |
| CoTBase LLM=Qwen2.5-VL-7B-Instruct2026.04 | 57.86 | 43 | |
| MM1.5-1B-MoEModel Scale=1B, Architecture=MoE, Zero-shot=true, Decoding Strategy=greedy2024.09 | 57.8 | — | |
| CoTBase LLM=Gemma3-12B-Instruct2026.04 | 57.69 | 49.92 | |
| R1-VLParameters=2B2026.06 | 57.6 | — | |
| ViGoRLParameters=3B, Setting=UGround2026.06 | 57.5 | — | |
| InternVL2-2BModel Scale=1B, Zero-shot=true, Decoding Strategy=beam search2024.09 | 57.4 | — | |
| DirectBase LLM=Gemma3-12B-Instruct2026.04 | 57.34 | 50.96 | |
| Dynamic-LLaVA-7B†Training-Free=false, Image Token=115 (-80%)2024.12 | 57 | — | |
| MM1.5-3BModel Scale=3B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 56.9 | — | |
| GPT-4VZero-shot=true, Decoding Strategy=greedy2024.09 | 56.5 | — | |
| MiniCPM-V 2.0-3BModel Scale=3B, Zero-shot=true, Decoding Strategy=beam search2024.09 | 55.8 | — | |
| MM1-3BModel Scale=3B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 55.8 | — | |
| MM1-7BModel Scale=7B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 55.7 | — | |
| LLaVAOne Vision-0.5BModel Scale=1B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 55.6 | — | |
| IVTPTraining-Free=false2024.12 | 54.6 | — | |
| DAM-3Bmask=12025.10 | 54.3 | — | |
| LLaVA-1.5-7BImage Token=576, TFLOPS=10.12024.12 | 53.7 | — | |
| LLaVA-FastVTraining-Free=true, Image Token=144 (-75%)2024.12 | 53.7 | — | |
| MM1.5-1BModel Scale=1B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 53.3 | — | |
| Reason-RFTParameters=2B2026.06 | 53.1 | — | |
| SINKTRACKBase LLM=Gemma3-4B-Instruct2026.04 | 52.46 | 47.39 | |
| Claude-3 Sonnetopen-source=false2024.04 | 51.9 | — | |
| DirectBase LLM=Gemma3-4B-Instruct2026.04 | 51.33 | 48.43 | |
| MM1-1BModel Scale=1B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 51.2 | — | |
| Claude-3 Opusopen-source=false2024.04 | 49.8 | — | |
| CoTBase LLM=Gemma3-4B-Instruct2026.04 | 49.41 | 44.53 | |
| SINKTRACKBase LLM=Qwen2.5-VL-3B-Instruct2026.04 | 48.28 | 28.42 | |
| DirectBase LLM=Qwen2.5-VL-7B-Instruct2026.04 | 47.49 | 39.93 | |
| LLaVA-1.5-7B+H2OTFLOPS=10.12024.12 | 42.3 | — | |
| CoTBase LLM=Qwen2.5-VL-3B-Instruct2026.04 | 35.69 | 26.46 | |
| DirectBase LLM=Qwen2.5-VL-3B-Instruct2026.04 | 31.15 | 23.25 | |
| PAM-3Bmask=12025.10 | 1.7 | — |