Instruction Following on ComplexBench
85.22Overall ScoreImpRIF-32B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ImpRIF-32BTraining Stage=SFT+RL2026.02 | 85.22 | — | |
| ImpRIF-32BTraining Stage=SFT2026.02 | 84.1 | — | |
| Qwen3-235B-A22BTraining Stage=Base2026.02 | 83.46 | — | |
| ImpRIF-32BTraining Stage=RL2026.02 | 83.39 | — | |
| ImpRIF-8BTraining Stage=SFT+RL2026.02 | 83.29 | — | |
| Qwen3-14BTraining Stage=Base2026.02 | 82.16 | — | |
| Qwen3-32BTraining Stage=Base2026.02 | 81.99 | — | |
| ImpRIF-8BTraining Stage=SFT2026.02 | 81.56 | — | |
| Qwen2.5-72BReasoning Capability=Lacks inference reasoning capabilities2026.02 | 81.48 | — | |
| Qwen3-8BTraining Stage=Base2026.02 | 81.37 | — | |
| ImpRIF-8BTraining Stage=RL2026.02 | 81.22 | — | |
| ImpRIF-4BTraining Stage=SFT+RL2026.02 | 80.91 | — | |
| ImpRIF-4BTraining Stage=SFT2026.02 | 79.37 | — | |
| ImpRIF-4BTraining Stage=RL2026.02 | 77.57 | — | |
| Qwen3-4BTraining Stage=Base2026.02 | 77.04 | — | |
| Ministral3-14BTraining Stage=Base2026.02 | 76.8 | — | |
| Llama3.1-70BReasoning Capability=Lacks inference reasoning capabilities2026.02 | 73.76 | — | |
| GPT-5.22026.06 | 71.6 | 86.7 | |
| Claude-4.6-Sonnet-Thinking2026.06 | 59.6 | 85.6 | |
| Gemini-3.1-Pro-Preview2026.06 | 56.2 | 87.6 | |
| Doubao-2.0-Pro2026.06 | 55.4 | 87.4 | |
| DeepSeek-V3.2Size=671B2026.06 | 54.8 | 83.9 | |
| S1-DeepResearchSize=32B, Train=SFT2026.06 | 54.2 | 83.1 | |
| Kimi-K2.5Size=1T2026.06 | 51.2 | 83.5 | |
| Qwen3.5-397BSize=397B2026.06 | 50.4 | 86.1 | |
| GLM-5Size=744B2026.06 | 47.9 | 85 | |
| MiniMax-M2.7Size=230B2026.06 | 45.1 | 79.2 | |
| Qwen3-235BSize=235B2026.06 | 43.6 | 79.8 | |
| Qwen3-32BSize=32B2026.06 | 40.6 | 77 | |
| MuSCBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 0.7 | — | |
| MuSCBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 0.6939 | — | |
| Self-Reward w/ BSMBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 0.6743 | — | |
| ISHEEPBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 0.6732 | — | |
| Qwen2-7B-InstructBase Model=Qwen2-7B-Instruct, Setting=Baseline2025.02 | 0.6724 | — | |
| ISHEEPBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 0.6713 | — | |
| Self-Reward w/ BSMBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 0.6702 | — | |
| Self-RewardBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 0.6698 | — | |
| Self-RewardBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 0.6645 | — | |
| MuSCBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 0.6598 | — | |
| SFTBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 0.6589 | — | |
| MuSCBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 0.6473 | — | |
| Self-CorrectBase Model=Qwen2-7B-Instruct, Setting=SelfInst2025.02 | 0.6441 | — | |
| Self-CorrectBase Model=Qwen2-7B-Instruct, Setting=PreInst2025.02 | 0.6432 | — | |
| Self-Reward w/ BSMBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 0.643 | — | |
| Self-Reward w/ BSMBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 0.6413 | — | |
| Self-Reward w/ GPT-4Base Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 0.6405 | — | |
| Self-Reward w/ GPT-4Base Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 0.6352 | — | |
| ISHEEPBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 0.6292 | — | |
| ISHEEPBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 0.6267 | — | |
| Self-RewardBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 0.6245 | — | |
| Self-RewardBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 0.6203 | — | |
| LLaMA-3-8B-InstructBase Model=LLaMA-3-8B-Instruct, Setting=Baseline2025.02 | 0.6149 | — | |
| Self-CorrectBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 0.6079 | — | |
| Self-CorrectBase Model=LLaMA-3-8B-Instruct, Setting=SelfInst2025.02 | 0.5591 | — | |
| SFTBase Model=LLaMA-3-8B-Instruct, Setting=PreInst2025.02 | 0.5393 | — |