Audio Perception and Reasoning on MMAR (CAFE framework overall)
63.51Perception AccuracyMPAR2-7B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MPAR2-7B2026.02 | 63.51 | 23.14 | 7.74 | 30.59 | 60.32 | |
| Step-Audio-R1.1Category=Large Audio Reasoning Models2026.02 | 56.18 | 23.77 | 11.25 | 21.49 | 67.5 | |
| Gemini-2.5-FlashCategory=Commercial Models2026.02 | 54.81 | 19.27 | 15.96 | 33.24 | 66.3 | |
| Step-Audio-R1Category=Large Audio Reasoning Models2026.02 | 53.08 | 24.82 | 12.42 | 6.77 | 67.4 | |
| Omni-R1Category=Large Audio Language Models2026.02 | 51.21 | 28.79 | 12.9 | 42.65 | 62.1 | |
| GPT-4o-AudioCategory=Commercial Models2026.02 | 48.68 | 17.2 | 21.71 | 38.05 | 63.8 | |
| MiMo-AudioCategory=Large Audio Language Models2026.02 | 42.79 | 27.7 | 22.4 | 48.84 | 59.87 | |
| Audio-Flamingo-3Category=Large Audio Reasoning Models, Evaluation Judge=GPT-52026.02 | 41.39 | 26.03 | 23.18 | 52.87 | 56.4 | |
| Qwen2.5-Omni-7BCategory=Large Audio Language Models2026.02 | 31.74 | 29.61 | 28.69 | 57.37 | 55.2 | |
| Audio ReasonerCategory=Large Audio Reasoning Models2026.02 | 27.44 | 25.55 | 32.58 | 55.72 | 36.8 | |
| DeSTA2.5-AudioCategory=Large Audio Language Models, Evaluation Judge=GPT-52026.02 | 23.19 | 33.95 | 35.57 | 69.32 | 41.6 | |
| Phi-4-MultimodalCategory=Large Audio Language Models2026.02 | 20.43 | 36.75 | 30.32 | 73.41 | 39.8 | |
| Qwen2-Audio-InstructCategory=Large Audio Language Models2026.02 | 8.27 | 56.46 | 28.58 | 81.01 | 29.9 |