Instruction Following on IFEval (IFEval Score)
94.64IFEval ScoreGPT-5-2025-08-07
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5-2025-08-07Setting=high2026.05 | 94.64 | |
| JTMode=Think2026.05 | 94.64 | |
| TeleChatMode=Thinking2026.05 | 94.27 | |
| Kimi-K2.52026.05 | 93.9 | |
| GLM-5Precision=FP82026.05 | 93.16 | |
| Step-3.5-Flash2026.05 | 93.16 | |
| o4-mini-2025-04-16Setting=high2026.05 | 92.98 | |
| Gemini-3-Pro-Preview2026.05 | 92.79 | |
| GPT-5-mini-2025-08-07Setting=high2026.05 | 92.79 | |
| Kimi-K2Mode=Thinking2026.05 | 92.42 | |
| o3-2025-04-162026.05 | 92.24 | |
| o3-mini-2025-04-16Setting=high2026.05 | 91.87 | |
| DeepSeek-V3.2-Speciale2026.05 | 91.68 | |
| Qwen3.5-9B (released)Model Size=9B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 91.31 | |
| MiniMax-M2.52026.05 | 91.13 | |
| GLM-4-72026.05 | 90.2 | |
| gpt-oss-120bSetting=high2026.05 | 90.2 | |
| MiniMax-M22026.05 | 90.2 | |
| GPT-5-nano-2025-08-07Setting=high2026.05 | 90.2 | |
| LongCat-Flash-Chat2026.05 | 90.2 | |
| Gemini-2.5-Pro2026.05 | 90.02 | |
| Qwen3.5-4B (released)Model Size=4B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 90.02 | |
| DeepSeek-V3.22026.05 | 89.65 | |
| Qwen3Parameters=30B, Active Parameters=A3B, Mode=Thinking, Version=25072026.05 | 89.65 | |
| 8B-A-GRPOScale=8B2025.09 | 89.65 | |
| Mimo-V2-Flash2026.05 | 89.46 | |
| Qwen3-NextParameters=80B, Active Parameters=A3B, Mode=Thinking2026.05 | 89.46 | |
| gpt-oss-20bSetting=high2026.05 | 88.91 | |
| ERNIE-4.5Parameters=300B, Active Parameters=A47B, Type=PT2026.05 | 88.91 | |
| GLM-4.62026.05 | 88.72 | |
| Finix-P1-32BMode=Thinking2026.05 | 88.72 | |
| Qwen3-4BMode=Thinking, Version=25072026.05 | 88.54 | |
| Kimi-K2Type=Instruct2026.05 | 88.54 | |
| Claude Sonnet 4Mode=Thinking2026.05 | 88.35 | |
| Qwen3Parameters=235B, Active Parameters=A22B, Type=Instruct, Version=25072026.05 | 88.35 | |
| GPT-4.1-202504142026.05 | 88.17 | |
| Qwen3Parameters=235B, Active Parameters=A22B, Mode=Thinking, Version=25072026.05 | 87.8 | |
| Qwen3-NextParameters=80B, Active Parameters=A3B, Type=Instruct2026.05 | 87.62 | |
| 8B-GRPOScale=8B2025.09 | 87.62 | |
| Llama4-MaverickParameters=17B, Experts=128E, Type=Instruct2026.05 | 87.06 | |
| Qwen3Parameters=235B, Active Parameters=A22B, Mode=Thinking2026.05 | 85.77 | |
| BaselinePrecision=BF162026.07 | 85.7 | |
| GLM-4.52026.05 | 85.4 | |
| GLM-4-32B-04142026.05 | 85.21 | |
| AFM-k5984d3s (Ours)Optimization=QAD + Speculative Decoding + SWA2026.07 | 84.5 | |
| GLM-Z1-32B-04142026.05 | 84.29 | |
| GLM-4.5-Air2026.05 | 84.1 | |
| Qwen3Parameters=30B, Active Parameters=A3B, Type=Instruct, Version=25072026.05 | 83.92 | |
| Seed-OSS-36BType=Instruct2026.05 | 83.73 | |
| Qwen3-4BType=Instruct, Version=25072026.05 | 82.44 | |
| Qwen3-8BScale=8B2025.09 | 82.26 | |
| DeepSeek-Chat-V3-03242026.05 | 81.89 | |
| QwQ-32B2026.05 | 81.52 | |
| ERNIE-4.5Parameters=21B, Active Parameters=A3B, Mode=Thinking2026.05 | 81.15 | |
| Gemma-3-27B-it2026.05 | 80.96 | |
| DeepSeek-R1-05282026.05 | 80.04 | |
| Qwen2.5-14B-InstructBase Model=Qwen2.5-14B-Instruct, # Bits=BF16, Decoding Strategy=greedy2026.05 | 79.85 | |
| Qwen3.5-2B (released)Model Size=2B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 79.48 | |
| doubao-seed-1.6Mode=thinking, Version=2506152026.05 | 78.93 | |
| FlexRound+LFQBase Model=Qwen2.5-14B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 78 | |
| FlexRoundBase Model=Qwen2.5-14B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 77.82 | |
| FlexRound+LFQBase Model=Qwen2.5-14B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 77.08 | |
| MiniCPM-Sala2026.05 | 76.89 | |
| Qwen3.5-9B + AR-SFTModel Size=9B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 76.52 | |
| OmniQuant+LFQBase Model=Qwen2.5-14B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 75.42 | |
| OmniQuant+LFQBase Model=Qwen2.5-14B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 75.23 | |
| FlexRoundBase Model=Qwen2.5-14B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 75.05 | |
| OmniQuantBase Model=Qwen2.5-14B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 74.31 | |
| OmniQuantBase Model=Qwen2.5-14B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 73.94 | |
| FLARE-4BModel Size=4B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 73.2 | |
| Block-AP+LFQBase Model=Qwen2.5-14B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 72.46 | |
| Qwen3.5-4B + AR-SFTModel Size=4B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 72.46 | |
| Block-AP+LFQBase Model=Qwen2.5-14B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 72.27 | |
| Block-APBase Model=Qwen2.5-14B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 71.72 | |
| FlexRound+LFQBase Model=Qwen2.5-7B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 71.35 | |
| FLARE-9BModel Size=9B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 71.35 | |
| Qwen2.5-7B-InstructBase Model=Qwen2.5-7B-Instruct, # Bits=BF16, Decoding Strategy=greedy2026.05 | 70.79 | |
| Block-APBase Model=Qwen2.5-14B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 70.79 | |
| FlexRoundBase Model=Qwen2.5-7B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 69.5 | |
| OmniQuant+LFQBase Model=Qwen2.5-7B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 69.5 | |
| FLARE-2BModel Size=2B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 68.95 | |
| OmniQuant+LFQBase Model=Qwen2.5-7B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 68.58 | |
| OmniQuantBase Model=Qwen2.5-7B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 68.21 | |
| OmniQuantBase Model=Qwen2.5-7B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 68.21 | |
| Block-AP+LFQBase Model=Qwen2.5-7B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 68.02 | |
| FlexRound+LFQBase Model=Qwen2.5-7B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 67.84 | |
| Qwen3.5-2B + AR-SFTModel Size=2B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 67.65 | |
| Block-APBase Model=Qwen2.5-7B-Instruct, # Bits=W4, Decoding Strategy=greedy2026.05 | 66.73 | |
| FlexRoundBase Model=Qwen2.5-7B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 66.54 | |
| Block-AP+LFQBase Model=Qwen2.5-7B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 63.77 | |
| Block-APBase Model=Qwen2.5-7B-Instruct, # Bits=W3g128, Decoding Strategy=greedy2026.05 | 61 | |
| Hunyuan-A13BType=Instruct2026.05 | 60.26 |