Human Evaluation
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Human Evaluation Portuguese N=10 (pt)
97.5Win Rate
3
Human Evaluation Polish N=10
90Win Rate
3
Human Evaluation Korean N=10
95Win Rate
3
Human Evaluation Italian N=10
90Win Rate
3
Human Evaluation French N=10
65Win Rate
3
Human Evaluation English N=10
80Win Rate
3
Human Evaluation German N=10 (de)
95Win Rate
3
Human Evaluation Arabic N=10
77.5Win Rate
3
Human Evaluation Feedback Level
72.3Validity Rate
3
Human Evaluation Evil Players
3.78Contributed Success
3
Human Evaluation Toxic Prompts
2Biased Item Count
3
Human Evaluation win rate
50.31Quality Win Rate
3
Human Evaluation 48 instances (test)
58.3Preference Rate
3
Human evaluation 1.0 (random subset of 30 responses)
96ICS
3
Human evaluation data
47.7Fact Matching Acc
3
Human Evaluation Tutoring Dialogues (test)
42.2Proactivity Win
3
Human Evaluation Set (test)
0.65Win Rate
3
Human Evaluation Set
3.9Quality Score
3
Human Evaluation set en-te 1.0 (test)
2.79Gender Agreement
3
Human Evaluation en-ru 1.0 (test)
2.81Gender Agreement
3
Human Evaluation set en-it 1.0 (test)
2.77Gender Agreement Score
3
Human Evaluation en-es 1.0 (test)
2.66Gender Agreement
3
Human Evaluation en-de 1.0 (test)
2.84Gender Agreement
3
Human Evaluation en-cs 1.0 (test)
2.65Gender Agreement
3
Human Evaluation set en-am 1.0 (test)
2.85Gender Agreement
3