Multi-modal Question Answering on MMMU (val)
70.7AccuracyProprietary API SOTA (Hurst et al., 2024)
Evaluation Results
| Method | Links | |
|---|---|---|
| Proprietary API SOTA (Hurst et al., 2024)Model Type=Proprietary API, Evaluation Protocol=MCQ2025.01 | 70.7 | |
| Gemini-1.5-flash + UnACBackbone=Gemini-1.5-flash, Prompting Strategy=UnAC2026.05 | 60.9 | |
| GPT4-V + UnACBackbone=GPT4-V, Prompting Strategy=UnAC2026.05 | 60.7 | |
| GPT4-V + SKETCHPADBackbone=GPT4-V, Prompting Strategy=SKETCHPAD2026.05 | 59.7 | |
| GPT4-V + CCoTBackbone=GPT4-V, Prompting Strategy=CCoT2026.05 | 58.7 | |
| GPT4-VBackbone=GPT4-V, Prompting Strategy=Baseline2026.05 | 57.2 | |
| GPT4-V + SoMBackbone=GPT4-V, Prompting Strategy=SoM2026.05 | 57.2 | |
| Open-Source SOTA (Chen et al., 2024d)Model Type=Open-Source (<10B), Evaluation Protocol=MCQ2025.01 | 56.2 | |
| Gemini-1.5-flashBackbone=Gemini-1.5-flash, Prompting Strategy=Baseline2026.05 | 56.1 | |
| InternVL2.0-8B + UnACBackbone=InternVL2.0-8B, Prompting Strategy=UnAC2026.05 | 54.7 | |
| InternVL2.0-8BBackbone=InternVL2.0-8B, Prompting Strategy=Baseline2026.05 | 51.8 | |
| LLaVA-OneVision-7B + UnACBackbone=LLaVA-OneVision-7B, Prompting Strategy=UnAC2026.05 | 51 | |
| LLaVA-OneVision-7BBackbone=LLaVA-OneVision-7B, Prompting Strategy=Baseline2026.05 | 48.8 | |
| IXC-2.5-ChatModel Type=Open-Source (<10B), Evaluation Protocol=MCQ2025.01 | 44.1 | |
| IXC-2.5Model Type=Open-Source (<10B), Evaluation Protocol=MCQ2025.01 | 42.9 | |
| LLaVA-v1.6-7B + UnACBackbone=LLaVA-v1.6-7B, Prompting Strategy=UnAC2026.05 | 37.4 | |
| LLaVA-v1.6-7BBackbone=LLaVA-v1.6-7B, Prompting Strategy=Baseline2026.05 | 36.9 |