LLM Alignment
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
11.85GPT-4o Win Rate
7
Jun 1, 2026
84Truthfulness Index
7
Feb 26, 2026
0.891Truthfulness Index
7
Feb 26, 2026
4.24HV
6
May 21, 2026
1.195Loss
6
Apr 27, 2026
51.8Instruction Following Win Rate
6
Apr 23, 2026
79.93Win Rate
6
Feb 26, 2026
81.9Win Rate
5
May 8, 2026
82.3Helpfulness Score
5
May 8, 2026
68.56Win Rate (GPT-4o)
4
Jun 1, 2026
55Win-rate
4
Feb 26, 2026
0.58Win Rate
4
Feb 26, 2026
62Win Rate
4
Feb 26, 2026
54.7Average Score
3
May 22, 2026
3.574Loss
3
Apr 27, 2026
58Win Rate
2
Feb 26, 2026
59Win Rate
2
Feb 26, 2026