Role-playing
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
86.66LLM-as-a-Judge Score
91
Jun 15, 2026
4.043Overall Average Score
70
Jun 26, 2026
4.525Overall Score
47
May 26, 2026
88.82LLM-as-a-Judge Score
42
Jun 15, 2026
94.46GPT-4o Win Rate
28
Feb 26, 2026
4.444MC
28
Feb 26, 2026
-0.026Avg Score
18
Feb 26, 2026
-0.956Turn Composition
18
Feb 26, 2026
-0.8Deviation Score (Literature)
18
Feb 26, 2026
-0.016RP Score (German)
18
Feb 26, 2026
-0.034R-EMI
18
Feb 26, 2026
83.88ROUGE-L (Haruhi)
12
Feb 26, 2026
3.38Average Score
10
May 29, 2026
0.67KR
7
Feb 26, 2026
4.98Engagement (EA)
7
Feb 26, 2026
81.97Alpaca-P Score
5
Jun 15, 2026
4.38Anthropomorphism
5
Jun 15, 2026
21.21K-On! ROUGE-L
5
Feb 26, 2026
83.35Alpaca-P Score
4
Jun 15, 2026
2.749CC
4
Feb 26, 2026
602Win Count
4
Feb 26, 2026
36.4Win Rate (vs GPT-4)
4
Feb 26, 2026