Instruction Following on CF-Bench
71Instruction Success RateClaude-Opus-4.7
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Claude-Opus-4.7Stage=-2026.05 | 71 | — | — | |
| R1-0528-Qwen3-8BStage=Iter32026.05 | 69 | — | — | |
| QwQ-32BStage=-2026.05 | 68 | — | — | |
| R1-0528-Qwen3-8BStage=Iter12026.05 | 68 | — | — | |
| R1-0528-Qwen3-8BStage=Iter22026.05 | 68 | — | — | |
| R1-0528-Qwen3-8BStage=BASE2026.05 | 66 | — | — | |
| GPT-4oStage=-2026.05 | 65 | — | — | |
| Distill-Qwen-14BStage=Iter32026.05 | 60 | — | — | |
| Distill-Qwen-14BStage=Iter22026.05 | 59 | — | — | |
| Distill-Qwen-14BStage=Iter12026.05 | 58 | — | — | |
| Distill-Qwen-14BStage=BASE2026.05 | 55 | — | — | |
| Self-Supervised-7BStage=-2026.05 | 52 | — | — | |
| Qwen2.5-7B-InstructStage=Iter32026.05 | 51 | — | — | |
| Qwen2.5-7B-InstructStage=Iter12026.05 | 50 | — | — | |
| Qwen2.5-7B-InstructStage=Iter22026.05 | 49 | — | — | |
| Qwen2.5-7B-InstructStage=BASE2026.05 | 47 | — | — | |
| MuSCBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 44 | 79 | 55 | |
| RAIF-7BStage=-2026.05 | 43 | — | — | |
| MuSCBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 42 | 78 | 54 | |
| ISHEEPBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 41 | 77 | 52 | |
| VERIF-8BStage=-2026.05 | 41 | — | — | |
| ISHEEPBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 40 | 76 | 52 | |
| Self-RewardBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 38 | 75 | 50 | |
| Self-Reward w/ BSMBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 38 | 75 | 50 | |
| Self-RewardBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 37 | 75 | 49 | |
| Self-Reward w/ BSMBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 37 | 75 | 49 | |
| SPAR-8B-DPOStage=-2026.05 | 37 | — | — | |
| Qwen2-7B-InstructBase Model=Qwen2-7B-Instruct, Setting=Baseline2025.02 | 36 | 74 | 49 | |
| Llama-3.1-8B-InstructStage=Iter32026.05 | 36 | — | — | |
| SFTBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 35 | 72 | 46 | |
| Llama-3.1-8B-InstructStage=Iter22026.05 | 35 | — | — | |
| Llama-3.1-8B-InstructStage=BASE2026.05 | 34 | — | — | |
| Llama-3.1-8B-InstructStage=Iter12026.05 | 34 | — | — | |
| MuSCBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 32 | 70 | 44 | |
| MuSCBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 30 | 69 | 42 | |
| ISHEEPBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 29 | 60 | 40 | |
| Self-Reward w/ BSMBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 29 | 68 | 40 | |
| ISHEEPBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 29 | 67 | 40 | |
| Self-Reward w/ BSMBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 28 | 68 | 39 | |
| Self-CorrectBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 28 | 67 | 38 | |
| Self-CorrectBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 28 | 66 | 37 | |
| Self-RewardBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 27 | 66 | 37 | |
| Self-RewardBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 26 | 65 | 35 | |
| Self-Reward w/ GPT-4Base Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 26 | 66 | 37 | |
| Self-Reward w/ GPT-4Base Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 25 | 66 | 37 | |
| Crab-7B-DPOStage=-2026.05 | 25 | — | — | |
| Conifer-7B-DPOStage=-2026.05 | 25 | — | — | |
| LLaMA-3-8B-InstructBase Model=LLaMA-3-8B-Instruct, Setting=Baseline2025.02 | 24 | 64 | 34 | |
| RM-DistillerPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=HH-RLHF, Algorithm=GRPO2026.01 | 24 | 64 | 33 | |
| Qwen2.5-1.5B-InstructStage=Iter12026.05 | 24 | — | — | |
| Qwen2.5-1.5B-InstructStage=Iter32026.05 | 24 | — | — | |
| Self-CorrectBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 23 | 60 | 32 | |
| RM-DistillerPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=HH-RLHF, Algorithm=PPO2026.01 | 23 | 60 | 29 | |
| RM-DistillerPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=HH-RLHF, Algorithm=DAPO2026.01 | 23 | 63 | 32 | |
| RM-DistillerPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=ShareGPT, Algorithm=DAPO2026.01 | 23 | 63 | 32 | |
| Qwen2.5-1.5B-InstructStage=Iter22026.05 | 23 | — | — | |
| BT ClassifierPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=HH-RLHF, Algorithm=GRPO2026.01 | 22 | 61 | 30 | |
| BT ClassifierPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=HH-RLHF, Algorithm=DAPO2026.01 | 22 | 62 | 31 | |
| RM-DistillerPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=ShareGPT, Algorithm=GRPO2026.01 | 22 | 62 | 31 | |
| Qwen2.5-1.5B-InstructStage=BASE2026.05 | 22 | — | — | |
| BT ClassifierPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=ShareGPT, Algorithm=GRPO2026.01 | 21 | 61 | 30 | |
| BT ClassifierPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=ShareGPT, Algorithm=DAPO2026.01 | 21 | 60 | 29 | |
| Self-CorrectBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 20 | 52 | 27 | |
| SFTBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 20 | 56 | 26 | |
| RM-DistillerPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=ShareGPT, Algorithm=PPO2026.01 | 19 | 59 | 27 | |
| BT ClassifierPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=HH-RLHF, Algorithm=PPO2026.01 | 18 | 57 | 26 | |
| BT ClassifierPolicy=Llama-3-8B, Student RM=Qwen2.5-3B-Instruct, Teacher=GPT-4o, Dataset (Training)=ShareGPT, Algorithm=PPO2026.01 | 17 | 56 | 25 | |
| BaselinePolicy=Llama-3-8B, Dataset (Training)=UltraChat, Algorithm=SFT2026.01 | 15 | 51 | 22 |