ResearchTasksMultimodal Evaluation ConsistencyFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedMLLM-as-a-Judge, RichHF-18K, GenAI-BenchGPT-4o44.2Average Score22Apr 30, 2026MLLM-as-a-JudgeGPT-4o39.6CO Score22Apr 30, 2026