Multimodal Large Language Model Evaluation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
56.7Average Score (All)
22
Apr 14, 2026
2,315MME Score
22
Apr 24, 2026
71.56Accuracy
18
Jun 2, 2026
43.7Reasoning
5
Mar 16, 2026
74.94MME
4
Feb 26, 2026
190Existence Score
3
Feb 26, 2026