LLM Evaluation on AlpacaEval 2.0
51.32LC Win RateSpecEM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SpecEMModel type=Methods of Ensembling Base LLMs2024.12 | 51.32 | 54.52 | — | — | — | |
| UniTEModel type=Methods of Ensembling Base LLMs2024.12 | 49.2 | 41.04 | — | — | — | |
| GenFuseModel type=Methods of Ensembling Base LLMs2024.12 | 49.06 | 50.84 | — | — | — | |
| Mistral-24b-instruct-2501Model type=Base LLMs2024.12 | 48.46 | 44.27 | — | — | — | |
| MOAModel type=Methods of Ensembling Base LLMs2024.12 | 46.98 | 51.24 | — | — | — | |
| Qwen2.5-32b-instructModel type=Base LLMs2024.12 | 43.82 | 43.54 | — | — | — | |
| Qwen2-72b-instructModel type=Base LLMs2024.12 | 38.1 | — | — | — | — | |
| Llama3-70b-instructModel type=Base LLMs2024.12 | 34.4 | 29.39 | — | — | — | |
| DPO-PoP-randomBase Model=Llama-3.1-8b, Data=Synthetic2025.09 | 14.62 | — | 14.78 | — | 1,909 | |
| DPO-PoP-iterBase Model=Llama-3.1-8b, Data=Synthetic2025.09 | 12.89 | — | 13.42 | — | 2,004 | |
| DPO-margin-gtBase Model=Llama-3.1-8b, Data=Synthetic2025.09 | 11.23 | — | 11.3 | — | 1,825 | |
| DPO-margin-1Base Model=Llama-3.1-8b, Data=Synthetic2025.09 | 11.07 | — | 11.06 | — | 1,864 | |
| DPO-margin-gt-scaledBase Model=Llama-3.1-8b, Data=Synthetic2025.09 | 10.95 | — | 11.43 | — | 1,881 | |
| Vanilla-DPOBase Model=Llama-3.1-8b, Data=Synthetic2025.09 | 10.38 | — | 10.56 | — | 1,869 | |
| OursBackbone=LLaMA-7B, Selection Ratio=50%2026.03 | 7.7 | — | 2 | — | — | |
| Full DatasetBackbone=LLaMA-7B, Selection Ratio=100%2026.03 | 6.7 | — | 1.9 | — | — | |
| Adapted RLCR2026.06 | — | — | — | 86.2 | — | |
| Qwen2-7B2026.06 | — | — | — | 85.7 | — | |
| SEE2026.06 | — | — | — | 90.8 | — |