RoleBench
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
RoleBench
85.67LLM-as-a-Judge Score
44
RoleBench (test)
88.82LLM-as-a-Judge Score
42
RoleBench (test)
36.4RAW Score
10
RoleBench Instruction Generalization
57.6CUS Score
10
RoleBench English 1.0 (Role Generalization)
60.2CUS Score
7
RoleBench Chinese instruction generalization 1.0
53.7ROUGE-L (CUS)
7
RoleBench instruction generalization
55.8GPT-4 Win Rate
5
RoleBench Chinese (instruction generalization)
36.4Win Rate (vs GPT-4)
4
RoleBench English Role Generalization
64.5Win Rate (GPT-4)
4