Real-world Visual Question Answering on MME-RealWorld-Lite (MMERW)
57AccuracyGPT-5-0807
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5-0807Model Category=Proprietary and Open-Source MLLMs2025.11 | 57 | |
| GPT-4o-0513Model Category=Proprietary and Open-Source MLLMs2025.11 | 51.6 | |
| SpatialThinker-30BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 49.2 | |
| SSL4RL-7B (Mask)Category=SSL4RL-7B, SSL Task=Mask2025.10 | 49.03 | |
| SSL4RL-7B (Hard-Contrastive)Category=SSL4RL-7B, SSL Task=Hard-Contrastive2025.10 | 48.61 | |
| SpatialThinker-7BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 48.3 | |
| Qwen2.5-VL-7B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 47.4 | |
| Claude-4-Sonnet-0514Model Category=Proprietary and Open-Source MLLMs2025.11 | 46.9 | |
| Qwen2.5-VL-3B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 46.7 | |
| SpatialThinker-3BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 46.5 | |
| Qwen2.5-VL-7B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 46.3 | |
| Qwen3-VL-30BModel Category=Proprietary and Open-Source MLLMs2025.11 | 45.8 | |
| Qwen2.5-VL-7BCategory=Base2025.10 | 45.59 | |
| Claude-3.5-Sonnet-0620Model Category=Proprietary and Open-Source MLLMs2025.11 | 45.2 | |
| Qwen2.5-VL 7B + NoisyGRPOModel=Qwen2.5-VL 7B, Training Strategy=+ NoisyGRPO2025.10 | 44.6 | |
| VLAA-Thinker-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 44.6 | |
| Qwen2.5-VL-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 44.1 | |
| Qwen2.5-VL 3B + SFTModel=Qwen2.5-VL 3B, Training Strategy=+ SFT2025.10 | 44 | |
| Qwen2.5-VL 3B + NoisyGRPOModel=Qwen2.5-VL 3B, Training Strategy=+ NoisyGRPO2025.10 | 44 | |
| Qwen2.5-VL 7BModel=Qwen2.5-VL 7B, Training Strategy=Base2025.10 | 43.6 | |
| Qwen2.5-VL 7B + SFTModel=Qwen2.5-VL 7B, Training Strategy=+ SFT2025.10 | 43.5 | |
| Qwen2.5-VL-3B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 43 | |
| Qwen2.5-VL 7B + GRPOModel=Qwen2.5-VL 7B, Training Strategy=+ GRPO2025.10 | 42.3 | |
| Qwen2.5-VL 3BModel=Qwen2.5-VL 3B, Training Strategy=Base2025.10 | 42.1 | |
| Qwen2.5-VL-3BModel Category=Proprietary and Open-Source MLLMs2025.11 | 41.9 | |
| Qwen2.5-VL 3B + GRPOModel=Qwen2.5-VL 3B, Training Strategy=+ GRPO2025.10 | 40.8 | |
| PositionCategory=SSL4RL-3B2025.10 | 38.19 | |
| JigsawCategory=SSL4RL-3B2025.10 | 35.12 | |
| RotationCategory=SSL4RL-3B2025.10 | 34.18 | |
| Qwen2.5-VL-3BCategory=Base2025.10 | 32.41 | |
| ContrastiveCategory=SSL4RL-3B2025.10 | 30.17 |