Helpfulness on AdvGLUE
75.15AccuracyDR-IRL
Evaluation Results
| Method | Links | |
|---|---|---|
| DR-IRLBackbone=Qwen-2-7B-Instruct2025.03 | 75.15 | |
| STAIRBackbone=Qwen-2-7B-Instruct2025.03 | 74.13 | |
| DPOBackbone=Qwen-2-7B-Instruct2025.03 | 70.97 | |
| DR-IRLBackbone=Llama-3.1-8B-Instruct2025.03 | 70.71 | |
| STAIRBackbone=Llama-3.1-8B-Instruct2025.03 | 69.2 | |
| IRLBackbone=Qwen-2-7B-Instruct2025.03 | 69.1 | |
| IRLBackbone=Llama-3.1-8B-Instruct2025.03 | 68.27 | |
| GRPOBackbone=Qwen-2-7B-Instruct2025.03 | 67.43 | |
| GRPOBackbone=Llama-3.1-8B-Instruct2025.03 | 66.93 | |
| SFTBackbone=Qwen-2-7B-Instruct2025.03 | 66.9 | |
| BaseBackbone=Qwen-2-7B-Instruct2025.03 | 66.5 | |
| DPOBackbone=Llama-3.1-8B-Instruct2025.03 | 66.27 | |
| Self-RewardingBackbone=Qwen-2-7B-Instruct2025.03 | 66.13 | |
| SACPOBackbone=Llama-3.1-8B-Instruct2025.03 | 65.6 | |
| CoTBackbone=Qwen-2-7B-Instruct2025.03 | 65.6 | |
| SACPOBackbone=Qwen-2-7B-Instruct2025.03 | 64.1 | |
| Self-RewardingBackbone=Llama-3.1-8B-Instruct2025.03 | 59.1 | |
| CoTBackbone=Llama-3.1-8B-Instruct2025.03 | 58.4 | |
| BaseBackbone=Llama-3.1-8B-Instruct2025.03 | 58.33 | |
| SFTBackbone=Llama-3.1-8B-Instruct2025.03 | 57.53 |