Instruction Following on IFEval (Accuracy)
90.39Accuracy (IFEval)Self Consistency (Best on Validation)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Self Consistency (Best on Validation)Category=Other Methods2025.07 | 90.39 | — | |
| Llama-3.3-70B-InstructAccess=Open-source, Parameters=70B2025.07 | 90 | — | |
| Simple RouterCategory=Other Methods2025.07 | 90 | — | |
| Open-source Upper BoundCategory=Baselines2025.07 | 90 | — | |
| SMCSCategory=Ours2025.07 | 90 | — | |
| MoACategory=Other Methods2025.07 | 89.33 | — | |
| Symbolic-MoE*Category=Other Methods2025.07 | 89 | — | |
| Claude-3.7-SonnetAccess=Close-source2025.07 | 88 | — | |
| Qwen-2.5-72B-InstructAccess=Open-source, Parameters=72B2025.07 | 86.3 | — | |
| Qwen3-30B-A3BBase Model=Qwen3-30B-A3B2026.05 | 86.3 | — | |
| GPT-4.1Access=Close-source2025.07 | 86 | — | |
| Self-MoACategory=Other Methods2025.07 | 86 | — | |
| ZEDASFTBase Model=Qwen3-30B-A3B2026.05 | 85.2 | — | |
| Qwen2.5-InstructModel Size=14B, GPU Hour=∼1.8M2025.05 | 85.01 | — | |
| ZEDABase Model=Qwen3-30B-A3B2026.05 | 84.3 | — | |
| Qwen3-32BAccess=Open-source, Parameters=32B2025.07 | 83.7 | — | |
| GLM-Z1-32B-0414Access=Open-source, Parameters=32B2025.07 | 83 | — | |
| Llama-3.3-Nemotron-Super-49B-v1Access=Open-source, Parameters=49B2025.07 | 82.7 | — | |
| AdaMoEBase Model=Qwen3-30B-A3B2026.05 | 82.4 | — | |
| NETSFTBase Model=Qwen3-30B-A3B2026.05 | 82.4 | — | |
| Tulu 3 8BDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 82.4 | — | |
| GPT-4oAccess=Close-source2025.07 | 82.3 | — | |
| Mistral-SmallModel Size=24B, GPU Hour=∼1.6M2025.05 | 82.25 | — | |
| GPT-o3-miniAccess=Close-source2025.07 | 82 | — | |
| TeleChat2-35B-32KAccess=Open-source, Parameters=35B2025.07 | 82 | — | |
| QwQ-32BAccess=Open-source, Parameters=32B2025.07 | 81.7 | — | |
| NETSFT→OPDBase Model=Qwen3-30B-A3B2026.05 | 81.7 | — | |
| Gemma-3-27b-itAccess=Open-source, Parameters=27B2025.07 | 81 | — | |
| AMARISRubric Strategy=per-instance rubric2026.05 | 81 | — | |
| DashAttentionModel Size=8B, Context=short-context2026.05 | 80.7 | — | |
| Majority VotingCategory=Other Methods2025.07 | 80.67 | — | |
| AMARISRubric Strategy=global rubric2026.05 | 80.6 | — | |
| Llama-3.1-8B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 80.6 | — | |
| Rubric-ARMRL Training Approach=Rubric-ARM2026.05 | 80.4 | — | |
| Claude-3.5-SonnetAccess=Close-source2025.07 | 80.3 | — | |
| DeepSeek-R1-Distill-Llama-70BAccess=Open-source, Parameters=70B2025.07 | 80.3 | — | |
| Qwen2.5-Coder-32B-InstructAccess=Open-source, Parameters=32B2025.07 | 80.3 | — | |
| InfiGFusionModel Size=14B, GPU Hour=1952025.05 | 80.22 | — | |
| FullAttnModel Size=8B, Context=short-context2026.05 | 80 | — | |
| NSAModel Size=8B, Context=short-context2026.05 | 80 | — | |
| RubricHub (RuRL)RL Training Approach=RuRL2026.05 | 79.8 | — | |
| OpenRubricsRL Training Approach=DPO via Rubric-RM2026.05 | 79.5 | — | |
| MiniLogitModel Size=14B, GPU Hour=2202025.05 | 79.36 | — | |
| RuscaRLRL Training Approach=RuscaRL2026.05 | 79 | — | |
| InfLLMv2Model Size=8B, Context=short-context2026.05 | 79 | — | |
| FuseChatModel Size=14B, GPU Hour=6502025.05 | 78.9 | — | |
| Qwen2.5-32b-InstructAccess=Open-source, Parameters=32B2025.07 | 78.7 | — | |
| RLCFRL Training Approach=RLCF2026.05 | 78.6 | — | |
| FuseLLMModel Size=14B, GPU Hour=2252025.05 | 78.56 | — | |
| Pivot-SFTModel Size=14B, GPU Hour=1202025.05 | 77.7 | — | |
| Phi-4Model Size=14B, GPU Hour=∼1.0M2025.05 | 77.34 | — | |
| Naive (no RL)RL Training Approach=None2026.05 | 77.3 | — | |
| EXAONE-Deep-32BAccess=Open-source, Parameters=32B2025.07 | 76.3 | — | |
| InfiFusionModel Size=14B, GPU Hour=1602025.05 | 76.02 | — | |
| OLMo-2-7B-1124-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 75.6 | — | |
| Qwen2.5-CoderModel Size=14B, GPU Hour=∼1.8M2025.05 | 74.7 | — | |
| Qwen-2.5-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 74.7 | — | |
| HuatuoGPT-o1-72BAccess=Open-source, Parameters=72B2025.07 | 74 | — | |
| DeepSeek-R1-Distill-Qwen-32BAccess=Open-source, Parameters=32B2025.07 | 73.7 | — | |
| Dynamic SkippingBase Model=Qwen3-30B-A3B2026.05 | 70.4 | — | |
| Gemma-2-9B-itDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 69.9 | — | |
| ZEDABase Model=GLM-4.7-Flash2026.05 | 68.2 | — | |
| OLMo-2-7B-SFTDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 68 | — | |
| GLM-4.7-FlashBase Model=GLM-4.7-Flash2026.05 | 67.3 | — | |
| ZEDASFTBase Model=GLM-4.7-Flash2026.05 | 66.4 | — | |
| NETSFTBase Model=GLM-4.7-Flash2026.05 | 65.8 | — | |
| NETSFT→OPDBase Model=GLM-4.7-Flash2026.05 | 65.1 | — | |
| InternLM2.5-20B-ChatAccess=Open-source, Parameters=20B2025.07 | 64.7 | — | |
| AdaMoEBase Model=GLM-4.7-Flash2026.05 | 63.8 | — | |
| Dynamic SkippingBase Model=GLM-4.7-Flash2026.05 | 63.4 | — | |
| EvaByte-SFTDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 60.2 | — | |
| BaselineBase Model=LLaDA-1.52026.05 | 58.41 | — | |
| Elastic-dLLMBase Model=LLaDA-1.52026.05 | 58.04 | — | |
| IBTPOBackbone=Qwen3-14B-Base2026.05 | 57.9 | — | |
| DPadBase Model=LLaDA-1.52026.05 | 57.86 | — | |
| SparseDBase Model=LLaDA-1.52026.05 | 57.67 | — | |
| Ministral-8B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 56.4 | — | |
| IBROBackbone=Qwen3-14B-Base2026.05 | 54.3 | — | |
| Vanilla GRPOBackbone=Qwen3-14B-Base2026.05 | 53.9 | — | |
| TreeRLBackbone=Qwen3-14B-Base2026.05 | 53.4 | — | |
| Initial ModelBackbone=Qwen3-14B-Base2026.05 | 52.9 | — | |
| BaselineBase Model=LLaDA-8B-Instruct2026.05 | 52.68 | — | |
| GRPO w/ Entropy RegBackbone=Qwen3-14B-Base2026.05 | 52.5 | — | |
| SparseDBase Model=LLaDA-8B-Instruct2026.05 | 52.31 | — | |
| Elastic-dLLMBase Model=LLaDA-8B-Instruct2026.05 | 52.31 | — | |
| DPadBase Model=LLaDA-8B-Instruct2026.05 | 51.2 | — | |
| IBTPOMaximum truncation length=8K2026.05 | 46.6 | — | |
| IBTPOMaximum truncation length=4K2026.05 | 46.4 | — | |
| OLMoE-1B-7B-0924-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 46.2 | — | |
| Initial ModelMaximum truncation length=4K2026.05 | 43.5 | — | |
| FP16Model=Llama-3.2-1B-Instruct2025.11 | 43.44 | — | |
| Initial ModelMaximum truncation length=8K2026.05 | 43.3 | — | |
| TreeRLMaximum truncation length=8K2026.05 | 42.6 | — | |
| Vanilla GRPOMaximum truncation length=8K2026.05 | 42.4 | — | |
| Vanilla GRPOMaximum truncation length=4K2026.05 | 42.3 | — | |
| TreeRLMaximum truncation length=4K2026.05 | 42.2 | — | |
| Quant-OnlyModel=Llama-3.2-1B-Instruct2025.11 | 41.4 | — | |
| IntAttentionModel=Llama-3.2-1B-Instruct2025.11 | 39.74 | — | |
| OLMo-v1.7-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 39.2 | — | |
| MAP-Neo-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 35.9 | — |