Multilingual Multimodal Reasoning on MMMB, Multilingual MMBench, and MTVQA Combined
67.5Overall AccuracyLaV-CoT-3B
Evaluation Results
| Method | Links | |
|---|---|---|
| LaV-CoT-3Btraining=SFT + GRPO2025.09 | 67.5 | |
| Gemini-2.5-flash2025.09 | 67.4 | |
| GPT-4o-2024-05-132025.09 | 66.6 | |
| InternVL3.5-8B2025.09 | 64.9 | |
| InternVL3-8B2025.09 | 64.7 | |
| LaV-CoT-3Btraining=SFT2025.09 | 64.7 | |
| Qwen2.5-VL-7B2025.09 | 64.4 | |
| Eagle2-9B2025.09 | 62.8 | |
| Qwen2-VL-7B2025.09 | 61.6 | |
| PARROT-7B2025.09 | 59.7 | |
| InternVL3.5-2B2025.09 | 58 | |
| InternVL3-2B2025.09 | 57.4 | |
| Qwen2.5-VL-3B2025.09 | 57.3 | |
| LLaVA-OneVision-7B2025.09 | 53.7 | |
| Qwen2-VL-2B2025.09 | 52.8 | |
| GLM-4v-9B2025.09 | 46.2 | |
| DeepSeek-VL-7B2025.09 | 43.9 | |
| Monkey-9.8B2025.09 | 35 |