Helpfulness on SimpleQA
6.64AccuracyDR-IRL
Evaluation Results
| Method | Links | |
|---|---|---|
| DR-IRLBackbone=Llama-3.1-8B-Instruct2025.03 | 6.64 | |
| STAIRBackbone=Llama-3.1-8B-Instruct2025.03 | 6.38 | |
| IRLBackbone=Llama-3.1-8B-Instruct2025.03 | 5.71 | |
| SFTBackbone=Llama-3.1-8B-Instruct2025.03 | 4.72 | |
| GRPOBackbone=Llama-3.1-8B-Instruct2025.03 | 4.48 | |
| DR-IRLBackbone=Qwen-2-7B-Instruct2025.03 | 4.47 | |
| DPOBackbone=Llama-3.1-8B-Instruct2025.03 | 4.46 | |
| IRLBackbone=Qwen-2-7B-Instruct2025.03 | 4.21 | |
| CoTBackbone=Llama-3.1-8B-Instruct2025.03 | 4.09 | |
| STAIRBackbone=Qwen-2-7B-Instruct2025.03 | 4.07 | |
| GRPOBackbone=Qwen-2-7B-Instruct2025.03 | 3.98 | |
| BaseBackbone=Qwen-2-7B-Instruct2025.03 | 3.79 | |
| SFTBackbone=Qwen-2-7B-Instruct2025.03 | 3.47 | |
| Self-RewardingBackbone=Qwen-2-7B-Instruct2025.03 | 3.37 | |
| CoTBackbone=Qwen-2-7B-Instruct2025.03 | 3.03 | |
| Self-RewardingBackbone=Llama-3.1-8B-Instruct2025.03 | 2.7 | |
| DPOBackbone=Qwen-2-7B-Instruct2025.03 | 2.59 | |
| BaseBackbone=Llama-3.1-8B-Instruct2025.03 | 2.52 | |
| SACPOBackbone=Llama-3.1-8B-Instruct2025.03 | 0.74 | |
| SACPOBackbone=Qwen-2-7B-Instruct2025.03 | 0.62 |