HH-RLHF
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
HH-RLHF
1.09MD Rate
68
HH-RLHF
54.3Accuracy
56
HH-RLHF
89.5Win Rate
45
HH-RLHF (test)
87.4Win Rate
36
HH-RLHF (test)
89.42Helpfulness Win Rate
31
HH-RLHF
61.4Accuracy
30
HH-RLHF (test)
0.87Diversity
23
HH-RLHF
59Accuracy
22
HH-RLHF (test)
1.02Harm Score
21
Anthropic HH-RLHF helpful core250 (test)
18.93Reward Score
18
HH-RLHF (test)
88.94Average Score
16
HH-RLHF
8.75Gemini Score
16
HH-RLHF (test)
0.476RK
16
HH-RLHF 300 prompts
69.8Win/Tie Rate vs Vanilla (GPT-4o)
16
HH-RLHF
74Human Win Rate
16
HH-RLHF
64.7Win Rate
14
HH-RLHF (held-out)
78Win Rate
14
HH-RLHF
81.3Coverage
12
HH-RLHF helpful core250 (held-out evaluation)
20.155Reward Score
12
HH-RLHF (test)
98Percent batches with BWR > 0.50
12
HH-RLHF
154Estimated Score (EST)
12
HH-RLHF
53BWR
12
HH-RLHF
47.3Win Rate
12
HH-RLHF harmless (test)
83.33Win Rate
12
HH-RLHF
0.472Rank Correlation (RK)
11