Multi-image Understanding on MMT (val)
67.4AccuracyInternVL2-Llama3-76B
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL2-Llama3-76B2026.01 | 67.4 | |
| GPT-4V2026.01 | 64.3 | |
| Qwen2VL-7B2026.01 | 61.7 | |
| InternVL2-8B2026.01 | 57.9 | |
| LLaVA-OV-7B2026.01 | 56.6 | |
| attention-masking strategybackbone=LLaVA-OV-7B, attention masking=true2026.01 | 55.3 | |
| Qwen2VL-2B2026.01 | 51.9 | |
| InternVL2-2B2026.01 | 46.7 | |
| attention-masking strategybackbone=LLaVA-OV-0.5B, attention masking=true2026.01 | 45.9 | |
| procedural data-generation strategybackbone=LLaVA-OV-0.5B, attention masking=false2026.01 | 45.6 | |
| LLaVA-OV-0.5B2026.01 | 41.1 |