Human Evaluation
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Human Evaluation Elo (test)
1,634Elo Score
34
Human Evaluation
1,090Score
21
Human Evaluation
1,148Score
21
Human Evaluation
1,096Score
21
Human Evaluation 50 participants, 400 ratings (test)
4.84Mean Score
16
Human Evaluation (Scene Cut Videos)
81.54Music Quality Win Rate
14
Human Evaluation 10 book excerpts
5Simplification
12
Human Evaluation 10 book excerpts
5Naturalness
12
Human Evaluation 1-5 scale
4.4Coherence
10
Human Evaluation Total
85Win Ratio
10
Human Evaluation Debate
86.6EA
10
Human Evaluation Creative Stories LLaMA3.1-8B-Instruct (test)
7.57Creativity Score
9
Human Evaluation (sample of 100)
75Successful Sentences Count
8
Human Evaluation (sample of 100)
15Successful Sentences
8
Human Evaluation N=20 (test)
19Win Count
8
Human Evaluation 40 LLM-generated prompts 1.0 (test)
1,114.8Total ELO
8
Five-question human evaluation set
4.6Relevance
8
Human Evaluation 30 volunteers (test)
7,082Win Rate
8
Human evaluation
87Visual Quality
8
Human Evaluation Solution Simulation (test)
3.75Score
8
Human Evaluation (User Study)
4.78Quality Score
7
Human Evaluation Set (test)
4.14Relevance
7
Human Evaluation 100-sentence sample (test)
3.74Simplicity
7
Human Evaluation
3.5Consistency Score
6
Human Evaluation 1-10 scale (test)
8.7Coherence
6