Human Evaluation
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Human Evaluation
7.62Accuracy
6
Human Evaluation User Actions Dataset (test)
79Win Rate
6
Human Evaluation 51 participants (test)
3.38Style Score
6
Human Evaluation
57.6VQE
6
Human Evaluation Average 2025 (test)
2.74Avg Human Eval Score
6
Human Evaluation EN⇒ZH 2025 (test)
2.61Human Evaluation Score
6
Human Evaluation ZH⇒EN 2025 (test)
3.01Human Evaluation Score
6
Human Evaluation User Study
3.71Naturalness
6
Human Evaluation 50 generations (test)
0.49Detoxification Count
6
Human Evaluation 50 groups: 5 motion patterns and 10 subjects
82.8Text Alignment
6
Human Evaluation Overall
66Win Rate
6
Human Evaluation Snow
56.5User Preference Score (%)
5
Human Evaluation Rain
41.3User Preference
5
Human Evaluation Weather Synthesis 5-point Likert scale (test)
4.16Photo-realism Score
5
Human Evaluation 50 papers sampled (test)
1,130Actionability (Elo)
5
Human Evaluation 40 document-query pairs
4.4Relevance
5
Human Evaluation
4.42Coherence Score
5
Human Evaluation Sketch Comedy (test)
3.45Funniness
5
Human Evaluation
3.84Quality Score
5
Human Evaluation
41.08Overall Preference Score
5
Human evaluation 100-sample set
3.7Factual Alignment
5
Human Evaluation
7.5PPP
5
Human Evaluation
4.4Coherence
5
Human Evaluation (In-Distribution)
85Win Ratio (Overall)
5
Human Evaluation User Study (test)
5.33Plausibility
4