Instruction Following on AlpacaEval 2.0
95.87Win RateAttention-MoA
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Attention-MoAModel Category=MoA-based2026.01 | 95.87 | 91.15 | — | — | — | |
| MoAModel Category=MoA-based2026.01 | 93.09 | 88.56 | — | — | — | |
| GPT-5 ChatLength controlled=true, Judge=GPT-4.12026.03 | 88.5 | — | — | — | — | |
| DeepSeek-V3.1Model Category=Individual Model2026.01 | 84.02 | 68.83 | — | — | — | |
| GPT-5.2Length controlled=true, Judge=GPT-4.12026.03 | 83.5 | — | — | — | — | |
| GPT-4.1Length controlled=true, Judge=GPT-4.12026.03 | 83.4 | — | — | — | — | |
| Gemini-2.5-ProModel Category=Individual Model2026.01 | 83.02 | 65.74 | — | — | — | |
| MoAaggregator=GPT-4o2026.05 | 78.7 | 65.7 | — | — | — | |
| RMoAModel Category=MoA-based2026.01 | 78.48 | 78.2 | — | — | — | |
| Gaussian time samplingJudge=GPT-4o, Evaluation Protocol=Averaged over 5 random seeds2026.05 | 78.27 | — | — | — | — | |
| Qwen-MaxModel Category=Individual Model2026.01 | 77.22 | 64.68 | — | — | — | |
| RLLMRM=Qwen3-8B (prompted), RM Type=Generative, RM Size=8B, Training samples=non-verifiable WildChat, Evaluation mode=thinking mode2026.03 | 77.1 | 71.4 | — | — | — | |
| DPOPModel=Gemma-2-9b-it, Average Response Length=19702026.06 | 73.76 | 78.22 | — | — | — | |
| CoDIT-Gemma3Student Model=Llama-3.1-8B2026.04 | 72.85 | 52.73 | — | — | — | |
| RLHFRM=Skywork-Reward-V2-Llama-3.1-8B, RM Type=Scalar, RM Size=8B, Training samples=non-verifiable WildChat, Evaluation mode=thinking mode2026.03 | 72.3 | 68.6 | — | — | — | |
| GOLFModel Backbone=Qwen-3-8B2026.03 | 71.8 | 71.94 | — | — | — | |
| CoDIT-Gemma3Student Model=Qwen3-8B-Base2026.04 | 71.41 | 47.98 | — | — | — | |
| RLHFRM=Nexusflow/Athene-RM-8B, RM Type=Scalar, RM Size=8B, Training samples=non-verifiable WildChat, Evaluation mode=thinking mode2026.03 | 71.2 | 70.9 | — | — | — | |
| DPOModel=Gemma-2-9b-it, Average Response Length=19002026.06 | 70.4 | 66.9 | — | — | — | |
| Qwen 3 32BThinking enabled=false, Length controlled=true, Judge=GPT-4.12026.03 | 69.6 | — | — | — | — | |
| Critique-GRPOModel Backbone=Qwen-3-8B2026.03 | 68.2 | 69.82 | — | — | — | |
| GPT-5 NanoLength controlled=true, Judge=GPT-4.12026.03 | 68 | — | — | — | — | |
| DPO+SamSBackbone=Gemma2-Instruct v0.2 (9B)2025.06 | 67.1 | 70.8 | — | — | — | |
| R-DPOBackbone=Gemma2-Instruct v0.2 (9B)2025.06 | 66.9 | 68.3 | — | — | — | |
| DPOBackbone=Gemma2-Instruct v0.2 (9B)2025.06 | 66.9 | 70.4 | — | — | — | |
| Pairwise-GRPOModel Backbone=Qwen-3-8B2026.03 | 66.34 | 68.34 | — | — | — | |
| SimPOModel=Gemma-2-9b-it, Average Response Length=17542026.06 | 66.12 | 73.08 | — | — | — | |
| CoDIT-Qwen3-8BStudent Model=Qwen3-8B-Base2026.04 | 65.62 | 54.25 | — | — | — | |
| Rubric-as-RewardModel Backbone=Qwen-3-8B2026.03 | 65.34 | 68.88 | — | — | — | |
| AlphaDPOModel=Gemma-2-9b-it, Average Response Length=16822026.06 | 65.32 | 74.9 | — | — | — | |
| Qwen3-8BRM=–, RM Type=–, RM Size=–, Training samples=non-verifiable WildChat, Evaluation mode=thinking mode2026.03 | 65.1 | 63.1 | — | — | — | |
| Direct-LikertModel Backbone=Qwen-3-8B2026.03 | 64.84 | 61.06 | — | — | — | |
| DPO (50%)Backbone=Gemma2-Instruct v0.2 (9B)2025.06 | 63.5 | 66.1 | — | — | — | |
| RMoAModel=GPT-4o2025.05 | 63.29 | — | — | — | — | |
| RoAd1Base Model=LLaMA2-7B, #Params.=0.02%, Finetuning Data=10K cleaned Alpaca2024.08 | 62.64 | — | — | — | — | |
| RoAd1Base Model=LLaMA2-7B, #Params.=0.02%, Finetuning Data=UltraFeedback2024.08 | 62.6 | — | — | — | — | |
| WizardLM 8×22B†2026.05 | 62.3 | 51.3 | — | — | — | |
| Claude-4.5-SonnetModel Category=Individual Model2026.01 | 61.74 | 73.49 | — | — | — | |
| LoReFTBase Model=LLaMA2-7B, #Params.=0.03%, Finetuning Data=UltraFeedback [7]2024.08 | 61.68 | — | — | — | — | |
| LoRABase Model=LLaMA2-7B, #Params.=0.83%, Finetuning Data=10K cleaned Alpaca2024.08 | 61.55 | — | — | — | — | |
| MoAModel=GPT-4o2025.05 | 60.55 | — | — | — | — | |
| CoDIT-Qwen3-30BStudent Model=Qwen3-8B-Base2026.04 | 60.36 | 52.74 | — | — | — | |
| LoReFTBase Model=LLaMA2-7B, #Params.=0.03%, Finetuning Data=10K cleaned Alpaca2024.08 | 60.21 | — | — | — | — | |
| Olmo 3.1 32BLength controlled=true, Judge=GPT-4.12026.03 | 60.1 | — | — | — | — | |
| MoA2026.05 | 59.8 | 65.1 | — | — | — | |
| I-DPO + MaPPOModel=Qwen2.5-32B-Instruct2025.07 | 58.68 | — | — | — | — | |
| IPOBackbone=Gemma2-Instruct v0.2 (9B)2025.06 | 58.4 | 62.6 | — | — | — | |
| MMoA2026.05 | 58 | 61.5 | — | — | — | |
| GPT-4.1Model Category=Individual Model2026.01 | 57.23 | 69.83 | — | — | — | |
| MoA-Lite2026.05 | 57 | 59.3 | — | — | — | |
| SMoAModel=GPT-4o2025.05 | 56.24 | — | — | — | — | |
| Mutual-TaughtCategory=Our Methods, Iteration=22025.05 | 55.9 | 54.1 | 2,177 | — | — | |
| DPO + Topo-TPO2026.05 | 55.6 | — | — | — | — | |
| SP2DPO (best-by-LC config)Student Backbone=Gemma-3-4B-IT, Config Source=SA-Q-v32026.01 | 55.59 | 42.15 | — | — | — | |
| KTOBackbone=Gemma2-Instruct v0.2 (9B)2025.06 | 55.5 | 61.7 | — | — | — | |
| PROSPERBackbone=Qwen2.5-7B-Instruct, Model Scale=7B, LLM Judge=gpt-5-mini2026.02 | 55.4 | — | — | — | — | |
| DPO + TPO2026.05 | 55.4 | — | — | — | — | |
| PROSPER-JCBackbone=Qwen2.5-7B-Instruct, Model Scale=7B, LLM Judge=gpt-5-mini2026.02 | 55.3 | — | — | — | — | |
| CoDIT-Qwen3-8BStudent Model=Llama-3.1-8B2026.04 | 55.3 | 42.75 | — | — | — | |
| GPT-4oModel=GPT-4o2025.05 | 55.18 | — | — | — | — | |
| Qwen-3-8BModel Backbone=Qwen-3-8B2026.03 | 55.16 | 52.6 | — | — | — | |
| Base (instruction-tuned)Student Backbone=Gemma-3-4B-IT2026.01 | 54.97 | 38.96 | — | — | — | |
| SP2DPO (JMAMP)Student Backbone=Gemma-3-4B-IT, J=3, K=32026.01 | 54.97 | 41.02 | — | — | — | |
| HyPOBackbone=Mistral-Nemo-Instruct (12B)2026.02 | 54.9 | 55.7 | — | — | — | |
| CoDIT-Qwen3-30BStudent Model=Llama-3.1-8B2026.04 | 54.71 | 46.03 | — | — | — | |
| DPOStudent Backbone=Gemma-3-4B-IT, Beta Selection=Best β ∈ {0.1, 0.3, 0.5} by LC2026.01 | 54.22 | 41.08 | — | — | — | |
| Rand-βi (U[0.03, 0.3])Student Backbone=Gemma-3-4B-IT, Beta Distribution=U[0.03, 0.3]2026.01 | 54.1 | 39.35 | — | — | — | |
| Quant.npuBackbone=Llama-3.2-3B-Instruct, Quantization Precision=W8A8, Reference Model=FP162026.05 | 53.59 | 48.82 | 2,406 | — | — | |
| GOLFModel Backbone=Llama-3.1-8B-Instruct2026.03 | 53.42 | 69.67 | — | — | — | |
| CPOBackbone=Gemma2-Instruct v0.2 (9B)2025.06 | 53.4 | 56.4 | — | — | — | |
| Llama-3-8B-S-SPPO Iter3Judge=GPT-4 Turbo Annotator, Backbone=Llama-3-8B, Iteration=32026.06 | 52.19 | 47.46 | 2,287 | — | — | |
| DPO2026.05 | 52.1 | — | — | — | — | |
| DPO + MaPPOModel=Qwen2.5-32B-Instruct2025.07 | 51.68 | — | — | — | — | |
| GPT-4 Omnisnapshot=05/132026.05 | 51.3 | 57.5 | — | — | — | |
| PROSPER-VBBackbone=Qwen2.5-7B-Instruct, Model Scale=7B, LLM Judge=gpt-5-mini2026.02 | 51.2 | — | — | — | — | |
| vs. Qwen2.5-32BTraining Steps=5002026.02 | 51.18 | 35.55 | — | — | — | |
| Elo-EvolveTraining Steps=5002026.02 | 51.18 | 38.03 | — | — | — | |
| I-DPOModel=Qwen2.5-32B-Instruct2025.07 | 51.12 | — | — | — | — | |
| AvR Stage II RSFTInit Model=Llama-3-8B-Ins, Data scale=10k2025.06 | 51 | 42.5 | 2,687 | — | — | |
| AvR Stage I DPO +refine round 2Init Model=Llama-3-8B-Ins2025.06 | 50.8 | 35.5 | 2,963 | — | — | |
| GPT-4-1106-preview2024.05 | 50 | — | — | — | — | |
| GPT-4-Turbo*Base model=Notable baselines2025.07 | 50 | — | 2,049 | — | — | |
| GPT-4 Previewsnapshot=11/062026.05 | 50 | 50 | — | — | — | |
| GPT-4 Turbo2026.06 | 50 | 50 | — | — | — | |
| Qwen 3 8BThinking enabled=false, Length controlled=true, Judge=GPT-4.12026.03 | 49.7 | — | — | — | — | |
| AvR Stage I DPO +refine round 1Init Model=Llama-3-8B-Ins2025.06 | 49.2 | 39.1 | 2,562 | — | — | |
| DPOBackbone=Mistral-Nemo-Instruct (12B)2026.02 | 49.2 | 50.4 | — | — | — | |
| RLLMThinking Mode=Thinking, RM=Qwen3-1.7B, RM Type=Generative, RM Size=1.7B, Evaluator=GPT-4o2026.03 | 49.2 | 43.9 | — | — | — | |
| Point GRPOTraining Steps=5002026.02 | 49.01 | 37.41 | — | — | — | |
| AvR Stage II RSFT +length controlInit Model=Llama-3-8B-Ins, Data scale=4k (14k)2025.06 | 49 | 51.4 | 1,989 | — | — | |
| I-DPO +MaPPOBackbone=Qwen2.5-14B-Instruct2025.07 | 48.89 | — | — | — | — | |
| I-DPO + MaPPOModel=Qwen2.5-14B-Instruct2025.07 | 48.89 | — | — | — | — | |
| DPOCategory=Iterative Preference Optimization Methods, Iteration=32025.05 | 48.7 | 47.2 | 1,930 | — | — | |
| SPPOCategory=Iterative Preference Optimization Methods, Iteration=32025.05 | 48.5 | 46.4 | 2,128 | — | — | |
| Llama-3-8B-S-SPPO Iter2Judge=GPT-4 Turbo Annotator, Backbone=Llama-3-8B, Iteration=22026.06 | 48.47 | 44.77 | 2,208 | — | — | |
| vs. Qwen2.5-14BTraining Steps=5002026.02 | 48.2 | 35.84 | — | — | — | |
| DPO+SamSBackbone=Llama3-Instruct v0.2 (8B)2025.06 | 48.2 | 51.5 | — | — | — | |
| PPOBackbone=Qwen2.5-14B-Instruct, clip ratio=0.22025.07 | 48.13 | — | — | — | — | |
| Elo-EvolveTraining Steps=3002026.02 | 48.07 | 35.02 | — | — | — | |
| Point GRPOTraining Steps=3002026.02 | 47.76 | 33.23 | — | — | — |