Helpfulness on HHH (Accuracy)
90.71AccuracySTAIR
Evaluation Results
| Method | Links | |
|---|---|---|
| STAIRBackbone=Qwen-2-7B-Instruct2025.03 | 90.71 | |
| DR-IRLBackbone=Qwen-2-7B-Instruct2025.03 | 90.71 | |
| SFTBackbone=Qwen-2-7B-Instruct2025.03 | 89.74 | |
| SACPOBackbone=Qwen-2-7B-Instruct2025.03 | 89.6 | |
| Self-RewardingBackbone=Qwen-2-7B-Instruct2025.03 | 88.31 | |
| CoTBackbone=Qwen-2-7B-Instruct2025.03 | 88.3 | |
| DPOBackbone=Qwen-2-7B-Instruct2025.03 | 88.08 | |
| BaseBackbone=Qwen-2-7B-Instruct2025.03 | 87.87 | |
| IRLBackbone=Qwen-2-7B-Instruct2025.03 | 87.52 | |
| GRPOBackbone=Qwen-2-7B-Instruct2025.03 | 86.89 | |
| DR-IRLBackbone=Llama-3.1-8B-Instruct2025.03 | 86.16 | |
| STAIRBackbone=Llama-3.1-8B-Instruct2025.03 | 85.66 | |
| SACPOBackbone=Llama-3.1-8B-Instruct2025.03 | 85.21 | |
| IRLBackbone=Llama-3.1-8B-Instruct2025.03 | 85.13 | |
| GRPOBackbone=Llama-3.1-8B-Instruct2025.03 | 84.5 | |
| DPOBackbone=Llama-3.1-8B-Instruct2025.03 | 83.84 | |
| SFTBackbone=Llama-3.1-8B-Instruct2025.03 | 82.63 | |
| BaseBackbone=Llama-3.1-8B-Instruct2025.03 | 82.5 | |
| Self-RewardingBackbone=Llama-3.1-8B-Instruct2025.03 | 82.09 | |
| CoTBackbone=Llama-3.1-8B-Instruct2025.03 | 81.63 |