ResearchTasksGeneral multi-turn dialogue evaluationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedMT-Bench-101Qwen3-Max-Thinking93.62Normalized Score5Mar 25, 2026