Preference Modeling on Instruction Following
65.2AccuracyBTPO
Evaluation Results
| Method | Links | |
|---|---|---|
| BTPOBase Model=Llama3.1-8B-Instruct2025.10 | 65.2 | |
| BTBase Model=Llama3.1-8B-Instruct2025.10 | 63.4 | |
| BTPOBase Model=Llama3.2-3B-Instruct2025.10 | 61.4 | |
| BTPOBase Model=Qwen2.5-7B-Instruct2025.10 | 60.1 | |
| BTPOBase Model=Qwen2.5-3B-Instruct2025.10 | 58.7 | |
| BTBase Model=Qwen2.5-7B-Instruct2025.10 | 58.7 | |
| BTBase Model=Llama3.2-3B-Instruct2025.10 | 58.7 | |
| BTBase Model=Qwen2.5-3B-Instruct2025.10 | 57 | |
| GRPO (pair)Base Model=Llama3.1-8B-Instruct2025.10 | 53.4 | |
| GRAMBase Model=Qwen2.5-3B-Instruct2025.10 | 53.3 | |
| GRPO (pair)Base Model=Qwen2.5-7B-Instruct2025.10 | 52.2 | |
| GRPO (pair)Base Model=Qwen2.5-3B-Instruct2025.10 | 51.1 | |
| GRAMBase Model=Llama3.2-3B-Instruct2025.10 | 50.4 | |
| GRAMBase Model=Qwen2.5-7B-Instruct2025.10 | 50.2 | |
| GRPO (pair)Base Model=Llama3.2-3B-Instruct2025.10 | 50.1 | |
| GRPO (point)Base Model=Qwen2.5-3B-Instruct2025.10 | 50 | |
| GRPO (point)Base Model=Qwen2.5-7B-Instruct2025.10 | 50 | |
| GRPO (point)Base Model=Llama3.2-3B-Instruct2025.10 | 50 | |
| GRAMBase Model=Llama3.1-8B-Instruct2025.10 | 50 | |
| GRPO (point)Base Model=Llama3.1-8B-Instruct2025.10 | 50 |