Preference Alignment on AlpacaEval 2
4.66Win-rate DeltaPrefixMemory-Tuning
Evaluation Results
| Method | Links | |
|---|---|---|
| PrefixMemory-TuningAlignment Protocol=DPO, Number of samples=10K2025.06 | 4.66 | |
| LoRAAlignment Protocol=DPO, Number of samples=10K2025.06 | 3.52 | |
| PrefixMemory-TuningAlignment Protocol=SimPO, Number of samples=10K2025.06 | 1.74 | |
| LoRAAlignment Protocol=SimPO, Number of samples=10K2025.06 | 1.24 | |
| PrefixMemory-TuningAlignment Protocol=SFT, Number of samples=10K2025.06 | 0.76 | |
| LoRAAlignment Protocol=SFT, Number of samples=10K2025.06 | 0.49 |