ResearchDatasetsInternal BenchmarkFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsHuman Preference Predictioninternal benchmark82.2Accuracy14Mathematical ReasoningInternal Benchmark65.5Average Score5Agent Action Safety Verificationinternal benchmark 300-scenario95Verdict Accuracy5General Multimodal Intelligence EvaluationInternal Benchmark (test)61.6Overall Score5