Visual Question Answering on RealWorldQA (test)
79AccuracyGPT5 mini
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT5 miniModel Source Category=Closed-source VLMs, Evaluation Mode=Thinking Mode2026.01 | 79 | — | |
| Qwen3-VL 32BModel Source Category=Open-weight VLMs, Model Parameters=32B, Evaluation Mode=Thinking Mode2026.01 | 78.4 | — | |
| Qwen3-VL 30B-A3BModel Source Category=Open-weight VLMs, Model Parameters=30B-A3B, Evaluation Mode=Thinking Mode2026.01 | 77.4 | — | |
| Gemini-2.5 FlashModel Source Category=Closed-source VLMs, Evaluation Mode=Thinking Mode2026.01 | 76 | — | |
| MMFineReason-8BModel Source Category=Ours, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 75.6 | — | |
| GPT-4oInference mode=Zero-shot2024.10 | 75.4 | — | |
| MMFineReason-4BModel Source Category=Ours, Model Parameters=4B, Evaluation Mode=Thinking Mode2026.01 | 74.9 | — | |
| Qwen3-VL 8BModel Source Category=Open-weight VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 73.5 | — | |
| MMR1 8BModel Source Category=Open-source VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 71 | — | |
| SSL4RL-7B (Hard-Contrastive)Category=SSL4RL-7B, SSL Task=Hard-Contrastive2025.10 | 70.58 | — | |
| HoneyBee 8BModel Source Category=Open-source VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 70.5 | — | |
| Qwen2 VLInference mode=Zero-shot, Parameters=7B2024.10 | 69.7 | — | |
| OMR 7BModel Source Category=Open-source VLMs, Model Parameters=7B, Evaluation Mode=Thinking Mode2026.01 | 69.4 | — | |
| SSL4RL-7B (Mask)Category=SSL4RL-7B, SSL Task=Mask2025.10 | 68.88 | — | |
| Qwen2.5VLSize=7B, Type=AR, Samples=>9M2025.12 | 68.5 | — | |
| LLaVA-OV-7BModel=LLaVA-OV-7B2026.01 | 68.4 | — | |
| LLaVA-OV-7B w/ CLIModel=LLaVA-OV-7B w/ CLI2026.01 | 68.4 | — | |
| MMFineReason-2BModel Source Category=Ours, Model Parameters=2B, Evaluation Mode=Thinking Mode2026.01 | 68.2 | — | |
| LLaVA-OV-7B w/ SLIModel=LLaVA-OV-7B w/ SLI2026.01 | 68.2 | — | |
| DiffusionVLSize=7B, Type=Diff., Samples=738K2025.12 | 68 | — | |
| InternVL-2-26BModel=InternVL-2-26B2026.01 | 67.2 | — | |
| GPT-4o-miniInference mode=Zero-shot2024.10 | 67.1 | — | |
| LLaVA-OVSize=7B, Type=AR, Samples=7.8M2025.12 | 66.3 | — | |
| Qwen2.5-VL-7BCategory=Base2025.10 | 65.88 | — | |
| Qwen2.5VLSize=3B, Type=AR, Samples=>9M2025.12 | 65.4 | — | |
| InternVL-2-8BModel=InternVL-2-8B2026.01 | 64.4 | — | |
| Cambrian-1Size=8B, Type=AR, Samples=-2025.12 | 64.2 | — | |
| MiniCPM-Llama3-V 2.5Inference mode=Zero-shot, Parameters=8B2024.10 | 63.5 | — | |
| LLaDA-VSize=8B, Type=Diff., Samples=16.5M2025.12 | 63.2 | — | |
| Baichuan-omniInference mode=Zero-shot, Parameters=7B2024.10 | 62.6 | — | |
| DiffusionVLSize=3B, Type=Diff., Samples=738K2025.12 | 61.6 | — | |
| LLaVA-OV-7B w/ DeepStackModel=LLaVA-OV-7B w/ DeepStack2026.01 | 61.4 | — | |
| VITAInference mode=Zero-shot, Parameters=8x7B2024.10 | 59 | — | |
| IXC-2.5-7BModel=IXC-2.5-7B2026.01 | 57.5 | — | |
| LLaVA-OV-0.5B w/ CLIModel=LLaVA-OV-0.5B w/ CLI2026.01 | 56.7 | — | |
| Fine-Tuning (Clean)Target Model=Llama2026.05 | 56.6 | 0 | |
| LLaVA-OV-0.5BModel=LLaVA-OV-0.5B2026.01 | 56 | — | |
| LLaVA-OV-0.5B w/ SLIModel=LLaVA-OV-0.5B w/ SLI2026.01 | 55.1 | — | |
| Zero-Shot (Clean)Target Model=InternVL2026.05 | 54.9 | 5.9 | |
| Zero-Shot (Clean)Target Model=LLaVA2026.05 | 53.6 | 3.9 | |
| LLaVA-OV-0.5B w/ DeepStackModel=LLaVA-OV-0.5B w/ DeepStack2026.01 | 52.9 | — | |
| MMGUARD-BPHTarget Model=Llama2026.05 | 52.9 | 3.7 | |
| MMGUARD-CRSTarget Model=Llama2026.05 | 51.7 | 4.9 | |
| Zero-Shot (Clean)Target Model=Llama2026.05 | 49.7 | 6.9 | |
| VILA-13BModel=VILA-13B2026.01 | 41.9 | — | |
| Zero-Shot (Clean)Target Model=Gemma2026.05 | 36.6 | 19.6 | |
| Zero-Shot (Clean)Target Model=GLM2026.05 | 27.5 | 39.8 |