General Instruction Following on Arena-Hard v2
85.9Scoreo3
Evaluation Results
| Method | Links | |
|---|---|---|
| o3Type=Closed-source LLM2026.02 | 85.9 | |
| Gemini-2.5-ProType=Closed-source LLM2026.02 | 79 | |
| Qwen3-30B-A3B-Instruct + More Query Rubricsbackbone=Qwen3-30B-A3B-Instruct, supervision=More Query Rubrics2026.02 | 74.6 | |
| Nemotron-3-Puzzle-75B-A9BModel=Nemotron-3-Puzzle-75B-A9B, Precision=FP82026.07 | 69.8 | |
| GenRM-R-Align-14BReward Model=GenRM-R-Align-14B2026.02 | 60.2 | |
| Claude-3.7-SonnetType=Closed-source LLM2026.02 | 59.8 | |
| GenRM-R-Align-8BReward Model=GenRM-R-Align-8B2026.02 | 59.5 | |
| Deepseek-R1Type=Open-source LLM2026.02 | 56.8 | |
| GenRM-RLVR-14BReward Model=GenRM-RLVR-14B2026.02 | 55.9 | |
| GenRM-RLVR-8BReward Model=GenRM-RLVR-8B2026.02 | 53.1 | |
| Qwen3-14B-as-GenRMReward Model=Qwen3-14B-as-GenRM2026.02 | 51 | |
| GPT-4.1Type=Closed-source LLM2026.02 | 50 | |
| Qwen3-235B-InstructType=Open-source LLM2026.02 | 46.7 | |
| Qwen3-8B-as-GenRMReward Model=Qwen3-8B-as-GenRM2026.02 | 46.6 | |
| Baichuan-M2-32BType=Specialized LLM2026.02 | 45.8 | |
| Qwen3-32BType=Open-source LLM2026.02 | 44.5 | |
| HuatuoGPT-o1-72BType=Specialized LLM2026.02 | 43.2 | |
| Qwen3-4B-Instruct + More Query Rubricsbackbone=Qwen3-4B-Instruct, supervision=More Query Rubrics2026.02 | 41.2 | |
| Qwen3-4B-Instruct + Doctor Rubricsbackbone=Qwen3-4B-Instruct, supervision=Doctor Rubrics2026.02 | 39.7 | |
| Qwen3-4B-Instruct + Principle Rubricsbackbone=Qwen3-4B-Instruct, supervision=Principle Rubrics2026.02 | 37 | |
| Qwen3-4B-Instruct + Draft Rubricsbackbone=Qwen3-4B-Instruct, supervision=Draft Rubrics2026.02 | 34.9 | |
| Qwen3-30B-A3B-InstructType=Our Method baseline2026.02 | 33.9 | |
| Qwen3-8BReward Model=Baseline2026.02 | 26.5 | |
| +RL (Skywork-Reward-V2-Llama-3.1-8B)Model=Qwen2.5-7B2025.07 | 18.5 | |
| +RL (Skywork-Reward-V2-Qwen3-4B)Model=Qwen2.5-7B2025.07 | 17.9 | |
| Instruct (official)Model=Qwen2.5-7B2025.07 | 17.1 | |
| +RL (Skywork-Reward-Gemma-2-27B-v0.2)Model=Qwen2.5-7B2025.07 | 15.5 | |
| Qwen3-4B-InstructType=Our Method baseline2026.02 | 15 | |
| +RL (Skywork-Reward-Llama-3-8B-v0.2)Model=Qwen2.5-7B2025.07 | 12.2 | |
| +SFTModel=Qwen2.5-7B2025.07 | 9.9 | |
| +RL (Skywork-Reward-V2-Llama-3.1-8B)Model=Llama-3.1-8B2025.07 | 6.3 | |
| +RL (Skywork-Reward-V2-Qwen3-4B)Model=Llama-3.1-8B2025.07 | 6 | |
| Instruct (official)Model=Llama-3.1-8B2025.07 | 5.8 | |
| BaseModel=Qwen2.5-7B2025.07 | 5.6 | |
| +RL (Skywork-Reward-Gemma-2-27B-v0.2)Model=Llama-3.1-8B2025.07 | 3.8 | |
| +SFTModel=Llama-3.1-8B2025.07 | 3.1 | |
| BaseModel=Llama-3.1-8B2025.07 | 2 | |
| +RL (Skywork-Reward-Llama-3-8B-v0.2)Model=Llama-3.1-8B2025.07 | 1.6 |