ResearchTasksHuman Consistency EvaluationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedGenAI-BenchVQAScore38.4Kendall's Tau-c16May 22, 2026RichHF-18KGemini-2.5-Pro33.9Kendall's Tau11Apr 30, 2026MLLM-as-a-JudgeLLaVA-Critic30.3CO Consistency Score11Apr 30, 2026Q-Reasoning (test)Proposed Human-Like Reasoning Framework (detailed)51.4ROUGE-1 Score6Feb 26, 2026