LLM Alignment on HH-RLHF 100K samples (test)
82.3Helpfulness ScoreHard-Pair-GRPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Hard-Pair-GRPOBase Model=LLaMA-2-7B-Chat2026.05 | 82.3 | 85.7 | |
| ORPOBase Model=LLaMA-2-7B-Chat2026.05 | 80.5 | 83.2 | |
| DPOBase Model=LLaMA-2-7B-Chat2026.05 | 80.1 | 82.8 | |
| Soft-Pair-GRPOBase Model=LLaMA-2-7B-Chat2026.05 | 79.5 | 82.1 | |
| Standard GRPOBase Model=LLaMA-2-7B-Chat2026.05 | 78.2 | 81.5 |