Instruction Following on Arena Hard v0.1
4.5Wrong Rate (%)DPO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DPOTraining Dataset=Helpsteer22026.02 | 4.5 | — | — | — | |
| SimPOTraining Dataset=Helpsteer22026.02 | 7 | — | — | — | |
| HyPOTraining Dataset=Helpsteer22026.02 | 9.3 | — | — | — | |
| w/o RLModel=Llama-3-8B-Instruct2026.06 | 19 | — | -1.7 | — | |
| DPOBackbone=Qwen-2.5-7B2026.02 | 28.8 | — | — | — | |
| RLOO + SmoothModel=Llama-3-8B-Instruct2026.06 | 30.1 | — | -1.6 | — | |
| GSPO + SmoothModel=Llama-3-8B-Instruct2026.06 | 30.4 | — | -1.7 | — | |
| RLOOModel=Llama-3-8B-Instruct2026.06 | 30.5 | — | -1.8 | — | |
| GRPO + SmoothModel=Llama-3-8B-Instruct2026.06 | 30.8 | — | -1.9 | — | |
| GRPOModel=Llama-3-8B-Instruct2026.06 | 31.1 | — | -1.8 | — | |
| GSPOModel=Llama-3-8B-Instruct2026.06 | 31.6 | — | -1.5 | — | |
| RLOO + GraphAEModel=Llama-3-8B-Instruct2026.06 | 32.7 | — | -1.9 | — | |
| PPOModel=Llama-3-8B-Instruct2026.06 | 33.2 | — | -1.7 | — | |
| SimPOBackbone=Qwen-2.5-7B2026.02 | 33.8 | — | — | — | |
| SimPOBackbone=Mistral-Nemo-Instruct (12B)2026.02 | 33.9 | — | — | — | |
| GRPO + GraphAEModel=Llama-3-8B-Instruct2026.06 | 33.9 | — | -1.6 | — | |
| GSPO + GraphAEModel=Llama-3-8B-Instruct2026.06 | 34.2 | — | -2.1 | — | |
| DPOBackbone=Mistral-Nemo-Instruct (12B)2026.02 | 35.5 | — | — | — | |
| HyPOBackbone=Qwen-2.5-7B2026.02 | 36.2 | — | — | — | |
| HyPOBackbone=Mistral-Nemo-Instruct (12B)2026.02 | 38.9 | — | — | — | |
| w/o RLModel=Qwen2.5-7B-Instruct2026.06 | 45.7 | — | -2.1 | — | |
| RLOOModel=Qwen2.5-7B-Instruct2026.06 | 50.8 | — | -1.9 | — | |
| RLOO + SmoothModel=Qwen2.5-7B-Instruct2026.06 | 50.8 | — | -1.7 | — | |
| GRPO + SmoothModel=Qwen2.5-7B-Instruct2026.06 | 51 | — | -2 | — | |
| GRPOModel=Qwen2.5-7B-Instruct2026.06 | 51.5 | — | -1.9 | — | |
| PPOModel=Qwen2.5-7B-Instruct2026.06 | 52.9 | — | -1.7 | — | |
| GSPO + SmoothModel=Qwen2.5-7B-Instruct2026.06 | 54.4 | — | -2.1 | — | |
| RLOO + GraphAEModel=Qwen2.5-7B-Instruct2026.06 | 55.2 | — | -1.8 | — | |
| GSPOModel=Qwen2.5-7B-Instruct2026.06 | 55.5 | — | -2.4 | — | |
| GRPO + GraphAEModel=Qwen2.5-7B-Instruct2026.06 | 57.8 | — | -1.9 | — | |
| GSPO + GraphAEModel=Qwen2.5-7B-Instruct2026.06 | 59.5 | — | -1.9 | — | |
| AvR Stage IIBase Model=Llama-3-8B-Instruct2025.06 | — | 34.5 | — | 3,144 | |
| DPOMethod Category=RL Method, Base Model=Llama-3-8B-Instruct2025.06 | — | 32.6 | — | — | |
| gpt-3.5-turbo-0125Base Model=gpt-3.5-turbo-01252025.06 | — | 23.3 | — | — | |
| gpt-4-0613Base Model=gpt-4-06132025.06 | — | 37.9 | — | — | |
| Llama-3-8B-Instruct (Seed)Base Model=Llama-3-8B-Instruct2025.06 | — | 20.6 | — | 2,485 | |
| Meta-Rewarding LLMIteration=1, Method Category=Meta-Rewarding LLM, Base Model=Llama-3-8B-Instruct2025.06 | — | 25.1 | — | 2,395 | |
| Meta-Rewarding LLMIteration=2, Method Category=Meta-Rewarding LLM, Base Model=Llama-3-8B-Instruct2025.06 | — | 27.4 | — | 2,416 | |
| Meta-Rewarding LLMIteration=3, Method Category=Meta-Rewarding LLM, Base Model=Llama-3-8B-Instruct2025.06 | — | 27.6 | — | 2,501 | |
| Meta-Rewarding LLMIteration=4, Method Category=Meta-Rewarding LLM, Base Model=Llama-3-8B-Instruct2025.06 | — | 29.1 | — | 2,422 | |
| ORPOMethod Category=RL Method, Base Model=Llama-3-8B-Instruct2025.06 | — | 25.8 | — | — | |
| R-DPOMethod Category=RL Method, Base Model=Llama-3-8B-Instruct2025.06 | — | 33.1 | — | — | |
| Self-Rewarding LLMIteration=1, Method Category=Self-Rewarding LLM, Base Model=Llama-3-8B-Instruct2025.06 | — | 23.2 | — | 2,438 | |
| Self-Rewarding LLMIteration=2, Method Category=Self-Rewarding LLM, Base Model=Llama-3-8B-Instruct2025.06 | — | 26.3 | — | 2,427 | |
| Self-Rewarding LLMIteration=3, Method Category=Self-Rewarding LLM, Base Model=Llama-3-8B-Instruct2025.06 | — | 28.2 | — | 2,413 | |
| Self-Rewarding LLMIteration=4, Method Category=Self-Rewarding LLM, Base Model=Llama-3-8B-Instruct2025.06 | — | 27.3 | — | 2,448 | |
| SimPOMethod Category=RL Method, Base Model=Llama-3-8B-Instruct2025.06 | — | 33.8 | — | — |