Instruction Following on AlpacaEval 2.0 (test)
67.45LC Win Rate (%)Offline+Humanline (G2-9B Completions)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Offline+Humanline (G2-9B Completions)Base Model=Gemma2-27B-Instruct, Completions=Gemma2-9B-Instruct (G2-9B), LR=2.5e-6, beta=0.12025.09 | 67.45 | 61.37 | — | — | — | 1.64 | — | |
| OnlineBase Model=Gemma2-27B-Instruct, LR=2.5e-6, beta=0.12025.09 | 66.49 | 74.22 | — | — | — | 1.48 | — | |
| Muon-8L/AdamW-32Reference Baseline=AdamW-32, Judge Model=GPT-4 Turbo, Fine-tuning=Yes2025.09 | 59.93 | 60.48 | — | — | — | — | — | |
| Muon-32Reference Baseline=AdamW-32, Judge Model=GPT-4 Turbo, Fine-tuning=Yes2025.09 | 59.17 | 59.67 | — | — | — | — | — | |
| Muon-8L/AdamW-8DReference Baseline=AdamW-32, Judge Model=GPT-4 Turbo, Fine-tuning=Yes2025.09 | 57.91 | 58.47 | — | — | — | — | — | |
| Offline (G2-9B Completions)Base Model=Gemma2-27B-Instruct, Completions=Gemma2-9B-Instruct (G2-9B), LR=2.5e-6, beta=0.12025.09 | 56.58 | 45.17 | — | — | — | 1.68 | — | |
| GPRLRM Size=8B, RM Type=GPM, Iter./Ep.=32026.05 | 56.51 | 48.33 | — | — | — | — | 1,600 | |
| Offline+Humanline (L3-8B Completions)Base Model=Gemma2-27B-Instruct, Completions=Llama3-8B-Instruct (L3-8B), LR=2.5e-6, beta=0.32025.09 | 56.27 | 44.49 | — | — | — | 1.67 | — | |
| GPRLRM Size=2B, RM Type=GPM, Iter./Ep.=32026.05 | 51.08 | 45.21 | — | — | — | — | 1,699 | |
| Muon-8L/AdamW-32Reference Baseline=Muon-32, Judge Model=GPT-4 Turbo, Fine-tuning=Yes2025.09 | 50.37 | 50.75 | — | — | — | — | — | |
| Offline (L3-8B Completions)Base Model=Gemma2-27B-Instruct, Completions=Llama3-8B-Instruct (L3-8B), LR=2.5e-6, beta=0.32025.09 | 48.59 | 32.36 | — | — | — | 1.58 | — | |
| GEMMA-2-9B-ITModel Size=9B, Type=Instruct2025.12 | 48.18 | 36.99 | — | — | — | — | — | |
| Muon-8L/AdamW-8DReference Baseline=Muon-32, Judge Model=GPT-4 Turbo, Fine-tuning=Yes2025.09 | 48.14 | 48.17 | — | — | — | — | — | |
| BaselineBase Model=Gemma2-27B-Instruct2025.09 | 45.9 | 35.45 | — | — | — | 1.54 | — | |
| SimPORM Size=–, RM Type=–, Iter./Ep.=–2026.05 | 44.7 | 40.5 | — | — | — | — | 1,825 | |
| STACKELBERGGDA-FOLLOWERRole=Follower2025.12 | 44.57 | 34.47 | — | — | — | — | — | |
| SPPORM Size=8B, RM Type=BT, Iter./Ep.=32026.05 | 42.55 | 40.92 | — | — | — | — | 1,948 | |
| GPORM Size=2B, RM Type=BT, Iter./Ep.=32026.05 | 42.21 | 44.2 | — | — | — | — | 2,151 | |
| GRPORM Size=8B, RM Type=BT, Iter./Ep.=32026.05 | 41.92 | 40.51 | — | — | — | — | 1,893 | |
| GPORM Size=8B, RM Type=BT, Iter./Ep.=32026.05 | 40.37 | 38.56 | — | — | — | — | 1,969 | |
| DPORM Size=–, RM Type=–, Iter./Ep.=–2026.05 | 40.3 | 37.9 | — | — | — | — | 1,837 | |
| SPPORM Size=2B, RM Type=BT, Iter./Ep.=32026.05 | 40.01 | 42.12 | — | — | — | — | 2,136 | |
| GRPORM Size=2B, RM Type=BT, Iter./Ep.=32026.05 | 39.87 | 38.21 | — | — | — | — | 1,925 | |
| SPPORM Size=8B, RM Type=GPM, Iter./Ep.=32026.05 | 39.45 | 41.64 | — | — | — | — | 2,385 | |
| Weak-to-strong search (Llama-3-70B-Instruct)Vocabulary Access=cross vocabulary, Backbone=Llama-3-70B-Instruct, Strategy=Chunk-level Beam Search, W=2, K=2, L=302024.05 | 39.09 | 39.81 | 4.068 | -4.583 | — | — | — | |
| GPORM Size=8B, RM Type=GPM, Iter./Ep.=32026.05 | 38.98 | 41.54 | — | — | — | — | 3,249 | |
| Weak-to-strong searchVocabulary Access=cross vocabulary, Base Model=Llama-3-70B-Instruct, W=2, K=2, L=302024.05 | 37.92 | 38.43 | 4.019 | -4.616 | — | — | — | |
| GPORM Size=2B, RM Type=GPM, Iter./Ep.=32026.05 | 37.74 | 48.25 | — | — | — | — | 2,582 | |
| BoNVocabulary Access=cross vocabulary, Base Model=Llama-3-70B-Instruct, N=42024.05 | 36.6 | 36.38 | 3.869 | -4.676 | — | — | — | |
| SPPORM Size=2B, RM Type=GPM, Iter./Ep.=32026.05 | 36.06 | 45.61 | — | — | — | — | 2,498 | |
| BoN (Llama-3-70B-Instruct)Vocabulary Access=cross vocabulary, Backbone=Llama-3-70B-Instruct, Strategy=Best-of-N, N=42024.05 | 35.96 | 36.43 | 3.876 | -4.668 | — | — | — | |
| STACKELBERGGDA-LEADERRole=Leader2025.12 | 35.04 | 25.59 | — | — | — | — | — | |
| BaseVocabulary Access=cross vocabulary, Base Model=Llama-3-70B-Instruct2024.05 | 34.42 | 32.18 | 3.833 | -4.674 | — | — | — | |
| Base (Llama-3-70B-Instruct)Vocabulary Access=cross vocabulary, Backbone=Llama-3-70B-Instruct, Strategy=Base2024.05 | 34.42 | 32.18 | 3.833 | -4.674 | — | — | — | |
| LLAMA-3.1-TULU-3-8B-DPOBase Model=Llama-3.1-8B, Fine-tuning=DPO2025.12 | 33.37 | 40.15 | — | — | — | — | — | |
| QWEN2.5-7B-INSTRUCTModel Size=7B, Type=Instruct2025.12 | 29.52 | 29.91 | — | — | — | — | — | |
| MixDPO (GPT-4o-mini)Base Model=LLaMA3-8B-Instruct, Reward Source=GPT-4o-mini2026.02 | 29.47 | 25.59 | — | — | — | — | — | |
| MixDPO (Orig. reward)Base Model=LLaMA3-8B-Instruct, Reward Source=Original reward scores2026.02 | 29.02 | 28.01 | — | — | — | — | — | |
| MS-SWIFTStudent Model=Qwen3-1.7B2026.03 | 28.4 | 27.86 | — | — | — | — | — | |
| KDFlowStudent Model=Qwen3-1.7B2026.03 | 28.23 | 28.32 | — | — | — | — | — | |
| FSDP Student + TeacherStudent Model=Qwen3-1.7B2026.03 | 28.18 | 28.2 | — | — | — | — | — | |
| Weak-to-strong searchVocabulary Access=cross vocabulary, Base Model=Llama-3-8B-Instruct, W=4, K=4, L=302024.05 | 27.17 | 27.43 | 3.407 | -4.862 | — | — | — | |
| Qwen3-1.7BDistillation=w/o KD2026.03 | 26.09 | 21.99 | — | — | — | — | — | |
| Weak-to-strong search (Llama-3-8B-Instruct)Vocabulary Access=cross vocabulary, Backbone=Llama-3-8B-Instruct, Strategy=Chunk-level Beam Search, W=4, K=4, L=302024.05 | 25.96 | 26.73 | 3.431 | -4.859 | — | — | — | |
| BoNVocabulary Access=cross vocabulary, Base Model=Llama-3-8B-Instruct, N=162024.05 | 25.35 | 24.32 | 3 | -5.07 | — | — | — | |
| SimPOBase Model=LLaMA3-8B-Instruct2026.02 | 24.68 | 22.73 | — | — | — | — | — | |
| LLAMA-3.1-8B-INSTRUCTModel Size=8B, Type=Instruct2025.12 | 24.66 | 26.69 | — | — | — | — | — | |
| BaseVocabulary Access=cross vocabulary, Base Model=Llama-3-8B-Instruct2024.05 | 22.92 | 22.57 | 2.682 | -5.156 | — | — | — | |
| Base (Llama-3-8B-Instruct)Vocabulary Access=cross vocabulary, Backbone=Llama-3-8B-Instruct, Strategy=Base2024.05 | 22.92 | 22.57 | 2.682 | -5.156 | — | — | — | |
| BoN (Llama-3-8B-Instruct)Vocabulary Access=cross vocabulary, Backbone=Llama-3-8B-Instruct, Strategy=Best-of-N, N=162024.05 | 22.42 | 22.54 | 3.039 | -5.02 | — | — | — | |
| Weak-to-strong searchVocabulary Access=black box, Base Model=gpt-3.5-turbo-instruct, W=2, K=2, L=1002024.05 | 20.07 | 12.61 | 1.212 | -6.391 | — | — | — | |
| Weak-to-strong search (gpt-3.5-turbo-instruct)Vocabulary Access=black box, Backbone=gpt-3.5-turbo-instruct, Strategy=Chunk-level Beam Search, W=2, K=2, L=1002024.05 | 19.8 | 13.23 | 1.285 | -6.295 | — | — | — | |
| BoNVocabulary Access=black box, Base Model=gpt-3.5-turbo-instruct, N=42024.05 | 19.59 | 12.51 | 1.017 | -6.455 | — | — | — | |
| DPOBase Model=LLaMA3-8B-Instruct2026.02 | 19.53 | 20.25 | — | — | — | — | — | |
| TAG-INSTRUCTBase Model=LLaMA3-8b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 19.5 | 19.21 | — | — | 1.3 | — | — | |
| Weak-to-strong searchVocabulary Access=same vocabulary, Base Model=Llama-2-70b-chat, W=2, K=2, L=302024.05 | 19.1 | 18.14 | 2.29 | -5.425 | — | — | — | |
| Weak-to-strong search (Llama-2-70b-chat)Vocabulary Access=same vocabulary, Backbone=Llama-2-70b-chat, Strategy=Chunk-level Beam Search, W=2, K=2, L=302024.05 | 19.04 | 18.15 | 2.3 | -5.438 | — | — | — | |
| BoN (gpt-3.5-turbo-instruct)Vocabulary Access=black box, Backbone=gpt-3.5-turbo-instruct, Strategy=Best-of-N, N=42024.05 | 18.6 | 13.15 | 1.202 | -6.327 | — | — | — | |
| BoNVocabulary Access=same vocabulary, Base Model=Llama-2-70b-chat, N=42024.05 | 17.57 | 16.91 | 2.061 | -5.576 | — | — | — | |
| BoN (Llama-2-70b-chat)Vocabulary Access=same vocabulary, Backbone=Llama-2-70b-chat, Strategy=Best-of-N, N=42024.05 | 16.73 | 15.99 | 2.145 | -5.515 | — | — | — | |
| EFT (Llama-2-70b-chat)Vocabulary Access=same vocabulary, Backbone=Llama-2-70b-chat, Strategy=EFT, beta=12024.05 | 16.58 | 16.85 | 2.37 | -5.381 | — | — | — | |
| BaseVocabulary Access=same vocabulary, Base Model=Llama-2-70b-chat2024.05 | 16.18 | 14.98 | 1.902 | -5.641 | — | — | — | |
| Base (Llama-2-70b-chat)Vocabulary Access=same vocabulary, Backbone=Llama-2-70b-chat, Strategy=Base2024.05 | 16.18 | 14.98 | 1.902 | -5.641 | — | — | — | |
| BaseVocabulary Access=black box, Base Model=gpt-3.5-turbo-instruct2024.05 | 16 | 10.58 | 0.771 | -6.556 | — | — | — | |
| Base (gpt-3.5-turbo-instruct)Vocabulary Access=black box, Backbone=gpt-3.5-turbo-instruct, Strategy=Base2024.05 | 16 | 10.58 | 0.771 | -6.556 | — | — | — | |
| EFTVocabulary Access=same vocabulary, Base Model=Llama-2-70b-chat, beta=0.252024.05 | 15.54 | 14.45 | 1.905 | -5.581 | — | — | — | |
| CodecLMBase Model=LLaMA3-8b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 14.73 | 12.71 | — | — | 1.1 | — | — | |
| TAG-INSTRUCTBase Model=LLaMA3.2-3b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 14.28 | 14.81 | — | — | 1.19 | — | — | |
| Weak-to-strong searchVocabulary Access=same vocabulary, Base Model=Llama-2-7b-chat, W=4, K=4, L=302024.05 | 13.65 | 14.11 | 2.234 | -5.424 | — | — | — | |
| zephyr-7b-beta (π*)Vocabulary Access=weak supervision2024.05 | 13.2 | 11 | 1.138 | -6.143 | — | — | — | |
| Weak-to-strong search (Llama-2-7b-chat)Vocabulary Access=same vocabulary, Backbone=Llama-2-7b-chat, Strategy=Chunk-level Beam Search, W=4, K=4, L=302024.05 | 13.16 | 14.2 | 2.115 | -5.451 | — | — | — | |
| BoNVocabulary Access=same vocabulary, Base Model=Llama-2-7b-chat, N=162024.05 | 13.11 | 13.23 | 1.65 | -5.738 | — | — | — | |
| Evol-InstructBase Model=LLaMA3-8b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 12.29 | 9.9 | — | — | 0.99 | — | — | |
| WizardLM-dataBase Model=LLaMA3-8b, Teacher Model=Ministral-Instruct-8b, #Convs=192K2025.05 | 11.68 | 5.69 | — | — | 0.75 | — | — | |
| BoN (Llama-2-7b-chat)Vocabulary Access=same vocabulary, Backbone=Llama-2-7b-chat, Strategy=Best-of-N, N=162024.05 | 11.6 | 11.67 | 1.536 | -5.721 | — | — | — | |
| Auto-Instruct-EvolBase Model=LLaMA3-8b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 11.49 | 10.19 | — | — | 1.01 | — | — | |
| CodecLMBase Model=LLaMA3.2-3b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 11.24 | 11.19 | — | — | 1.03 | — | — | |
| Tree-InstructBase Model=LLaMA3-8b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 10.89 | 10.47 | — | — | 1.01 | — | — | |
| BaseVocabulary Access=same vocabulary, Base Model=Llama-2-7b-chat2024.05 | 10.08 | 10.3 | 1.183 | -5.849 | — | — | — | |
| Base (Llama-2-7b-chat)Vocabulary Access=same vocabulary, Backbone=Llama-2-7b-chat, Strategy=Base2024.05 | 10.08 | 10.3 | 1.183 | -5.849 | — | — | — | |
| EFT (Llama-2-7b-chat)Vocabulary Access=same vocabulary, Backbone=Llama-2-7b-chat, Strategy=EFT, beta=12024.05 | 10.07 | 11.63 | 1.924 | -5.535 | — | — | — | |
| Tree-InstructBase Model=LLaMA3.2-3b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 10.01 | 8.5 | — | — | 0.91 | — | — | |
| tulu-2-dpo-7bVocabulary Access=weak supervision, Backbone=Tulu-2-7b, Training=DPO2024.05 | 9.46 | 8.1 | 0.743 | -6.31 | — | — | — | |
| Alpaca-5kBase Model=LLaMA3-8b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 9.33 | 5.09 | — | — | 0.74 | — | — | |
| Evol-InstructBase Model=LLaMA3.2-3b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 9.06 | 7.48 | — | — | 0.85 | — | — | |
| EFTVocabulary Access=same vocabulary, Base Model=Llama-2-7b-chat, beta=0.252024.05 | 9.05 | 10.08 | 1.244 | -5.819 | — | — | — | |
| tulu-2-7bVocabulary Access=weak supervision, Backbone=Tulu-2-7b, Training=Reference model2024.05 | 9.03 | 5.38 | -1.07 | -7.362 | — | — | — | |
| LLAMA-3.1-TULU-3-8B-SFTBase Model=Llama-3.1-8B, Fine-tuning=SFT2025.12 | 8.83 | 14.26 | — | — | — | — | — | |
| Auto-Instruct-EvolBase Model=LLaMA3.2-3b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 8.79 | 8.79 | — | — | 0.95 | — | — | |
| Alpaca-5kBase Model=LLaMA3.2-3b, Teacher Model=Ministral-Instruct-8b, #Convs=5K2025.05 | 8.52 | 4.73 | — | — | 0.71 | — | — | |
| Alpaca-CleanBase Model=LLaMA3.2-3b, Teacher Model=Ministral-Instruct-8b, #Convs=52K2025.05 | 8.06 | 4.12 | — | — | 0.67 | — | — | |
| Alpaca-CleanBase Model=LLaMA3-8b, Teacher Model=Ministral-Instruct-8b, #Convs=52K2025.05 | 7.73 | 5.21 | — | — | 0.74 | — | — | |
| mistral-7b-sft-beta (πref)Vocabulary Access=weak supervision2024.05 | 7.54 | 4.77 | -1.274 | -7.618 | — | — | — | |
| WizardLM-dataBase Model=LLaMA3.2-3b, Teacher Model=Ministral-Instruct-8b, #Convs=192K2025.05 | 6.74 | 4.46 | — | — | 0.68 | — | — | |
| SelectiveDPOBase Model=LLaMA3-8B-Instruct2026.02 | 3.25 | 1.12 | — | — | — | — | — |