Instruction Following on AlpacaEval 2.0 (Pairwise Win Rates)
92.8Instruct vs. Pretrained Win RateCME-GRPO
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CME-GRPOGenerator=Llama-3.2-1B-It, Pretrained baseline=Llama-3.2-1B, Instruct baseline=Llama-3.2-1B-It, Training Regime=Instruction-Tuned baselines, Reference Model=gemma-3-1b-it, LLM Judge=GPT-5.2 and Claude Sonnet 4.62026.05 | 92.8 | 94.8 | 53.3 | 7.5 | 44.2 | 42.2 | |
| CME-GRPOGenerator=Qwen2.5-0.5B-It, Pretrained baseline=Qwen2.5-0.5B, Instruct baseline=Qwen2.5-0.5B-It, Training Regime=Instruction-Tuned baselines, Reference Model=Llama-3.2-1B-Instruct, LLM Judge=GPT-5.2 and Claude Sonnet 4.62026.05 | 77 | 83.5 | 52.5 | 7.8 | 21.7 | 20.2 | |
| CME-GRPOGenerator=OLMo-1B-SFT, Pretrained baseline=OLMo-1B-SFT, Instruct baseline=OLMo-1B-DPO, Training Regime=SFT baselines, Reference Model=gemma-3-1b-it, LLM Judge=GPT-5.2 and Claude Sonnet 4.62026.05 | 40.8 | 61.4 | 44 | 31 | 36.2 | 37.6 | |
| CME-GRPOGenerator=Qwen2.5-0.5B, Pretrained baseline=Qwen2.5-0.5B, Instruct baseline=Qwen2.5-0.5B-It, Training Regime=Pretrained baselines, Reference Model=gemma-3-1b-it, LLM Judge=GPT-5.2 and Claude Sonnet 4.62026.05 | 23 | 71.4 | 34 | 7 | 13.2 | 22.6 | |
| CME-GRPOGenerator=Llama-3.2-1B, Pretrained baseline=Llama-3.2-1B, Instruct baseline=Llama-3.2-1B-It, Training Regime=Pretrained baselines, Reference Model=gemma-3-1b-it, LLM Judge=GPT-5.2 and Claude Sonnet 4.62026.05 | 7.2 | 55 | 6.5 | 6.5 | 9.8 | 42.2 | |
| CME-GRPOGenerator=Gemma-3-1B-pt, Pretrained baseline=gemma-3-1b-pt, Instruct baseline=gemma-3-1b-it, Training Regime=Pretrained baselines, Reference Model=Llama-3.2-1B-Instruct, LLM Judge=GPT-5.2 and Claude Sonnet 4.62026.05 | 4 | 61.2 | 9.2 | 5.2 | 9.2 | 59.2 |