Instruction Following on IFBench (Score %)
82IFBench ScoreDS-v4-Flash
Evaluation Results
| Method | Links | |
|---|---|---|
| DS-v4-FlashPrompting protocol=prompt loose, Parameter Count=284B-A13B2026.06 | 82 | |
| Nemotron 3 UltraPrompting protocol=prompt loose, Parameter Count=550B-A55B2026.06 | 81.7 | |
| DS-v4-ProPrompting protocol=prompt loose, Parameter Count=1.6T-A49B2026.06 | 79.1 | |
| Qwen-3.5Prompting protocol=prompt loose, Parameter Count=397B-17B2026.06 | 78.2 | |
| GLM-5.1Prompting protocol=prompt loose, Parameter Count=744B-A40B2026.06 | 76.6 | |
| DeepSeek-V4-ProThinking mode=true2026.05 | 76 | |
| MiniMax-2.7Prompting protocol=prompt loose, Parameter Count=230B-A10B2026.06 | 74.6 | |
| GLM5.1Thinking mode=true2026.05 | 74 | |
| Kimi-K2.6Prompting protocol=prompt loose, Parameter Count=1T-A32B2026.06 | 73.7 | |
| Gemini 3.1 ProThinking mode=true2026.05 | 71.33 | |
| Kimi K2.6Thinking mode=true2026.05 | 68.67 | |
| GPT-5.5Thinking mode=true2026.05 | 67.33 | |
| Qwen3.5-397B-A17BThinking mode=true2026.05 | 65.67 | |
| TeacherDistillation Setting=Multi-Domain Distillation, Teacher Model=Qwen3-Nemotron-4B2026.05 | 62.93 | |
| Qwen3.5-4BActive / Total Parameters=4.0B / 4.0B, Sampling Settings=Recommended sampling settings2026.05 | 59.2 | |
| Ling-2.6-1T2026.06 | 57.62 | |
| GLM-5Thinking Mode=non-thinking2026.06 | 57.14 | |
| Qwen3-4B-Thinking-2507Active / Total Parameters=4.0B / 4.0B, Sampling Settings=Recommended sampling settings2026.05 | 52.9 | |
| ZAYA1-8BActive / Total Parameters=0.7B / 8.0B, Sampling Settings=T=1.0, top-p=0.95, top-k disabled2026.05 | 52.6 | |
| Hy-MT2-30B-A3BThinking mode=false2026.05 | 50.67 | |
| GLM5.1Thinking mode=false2026.05 | 50 | |
| DeepSeek-V3.2Thinking Mode=nothink2026.06 | 50 | |
| GPT-5.4Reasoning Mode=non-reasoning2026.06 | 49.73 | |
| Qwen3.5-397B-A17BThinking mode=false2026.05 | 48.33 | |
| Gemma4-26B-A4BThinking mode=false2026.05 | 45.6 | |
| GPT-5.5Thinking mode=false2026.05 | 43.33 | |
| STEP3-VL-10BNumber of Parameters=10B2026.01 | 43.28 | |
| Gemma-4-E4B-itActive / Total Parameters=4.0B / 8.0B, Sampling Settings=Recommended sampling settings2026.05 | 42.7 | |
| DeepSeek-V4-ProThinking mode=false2026.05 | 42.67 | |
| Kimi-K2.5Mode=Instant2026.06 | 42.53 | |
| TrOPDDistillation Setting=Multi-Domain Distillation, Student Model=Qwen3-SFT-1.7B, Teacher Model=Qwen3-Nemotron-4B2026.05 | 42.18 | |
| HeRLBackbone=Qwen2.5-7B-Instruct, Training Method=HeRL2026.03 | 39.7 | |
| HeRLBackbone=Qwen3-4B-Instruct-2507, Training Method=HeRL2026.03 | 39.7 | |
| Entropy OPDDistillation Setting=Multi-Domain Distillation, Student Model=Qwen3-SFT-1.7B, Teacher Model=Qwen3-Nemotron-4B2026.05 | 38.78 | |
| Kimi K2.6Thinking mode=false2026.05 | 38 | |
| OPDDistillation Setting=Multi-Domain Distillation, Student Model=Qwen3-SFT-1.7B, Teacher Model=Qwen3-Nemotron-4B2026.05 | 37.07 | |
| RLVRBackbone=Qwen3-4B-Instruct-2507, Training Method=RLVR2026.03 | 36.9 | |
| Qwen3-VL ThinkingNumber of Parameters=8B2026.01 | 36.65 | |
| EOPDDistillation Setting=Multi-Domain Distillation, Student Model=Qwen3-SFT-1.7B, Teacher Model=Qwen3-Nemotron-4B2026.05 | 36.39 | |
| REOPOLDDistillation Setting=Multi-Domain Distillation, Student Model=Qwen3-SFT-1.7B, Teacher Model=Qwen3-Nemotron-4B2026.05 | 36.05 | |
| Qwen3.6-35B-A3BThinking mode=false2026.05 | 36 | |
| Hy-MT2-1.8BThinking mode=false2026.05 | 35.33 | |
| Hy-MT2-7BThinking mode=false2026.05 | 35.33 | |
| Gemma4-E4BThinking mode=false2026.05 | 32 | |
| RLVRBackbone=Qwen2.5-7B-Instruct, Training Method=RLVR2026.03 | 31.6 | |
| SFTBackbone=Qwen3-4B-Instruct-2507, Training Method=SFT2026.03 | 31.3 | |
| HeRLBackbone=Llama-3.2-3B-Instruct, Training Method=HeRL2026.03 | 30.6 | |
| Qwen3-4B-Instruct-2507Backbone=Qwen3-4B-Instruct-2507, Training Method=Initial2026.03 | 29.9 | |
| InternVL 3.5Number of Parameters=8B2026.01 | 28.14 | |
| SFTBackbone=Qwen2.5-7B-Instruct, Training Method=SFT2026.03 | 27.9 | |
| DPOBackbone=Qwen3-4B-Instruct-2507, Training Method=DPO2026.03 | 27.9 | |
| GLM-4.6V FlashNumber of Parameters=9B2026.01 | 27.47 | |
| RLVRBackbone=Llama-3.2-3B-Instruct, Training Method=RLVR2026.03 | 26.6 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct, Training Method=Initial2026.03 | 26.2 | |
| Qwen3-SFT-1.7BDistillation Setting=Baseline (None), Student Model=Qwen3-SFT-1.7B2026.05 | 26.19 | |
| Gemma4-E2BThinking mode=false2026.05 | 26 | |
| DPOBackbone=Qwen2.5-7B-Instruct, Training Method=DPO2026.03 | 25.9 | |
| SFTBackbone=Llama-3.2-3B-Instruct, Training Method=SFT2026.03 | 24.8 | |
| Llama-3.2-3B-InstructBackbone=Llama-3.2-3B-Instruct, Training Method=Initial2026.03 | 23.8 | |
| MiMo-VL RL-2508Number of Parameters=7B2026.01 | 23.47 | |
| DPOBackbone=Llama-3.2-3B-Instruct, Training Method=DPO2026.03 | 22.1 | |
| Qwen3-1.7B# Total Params=1.7B, # Trained Tokens=36T2025.11 | 21.27 | |
| LFM2-1.2B# Total Params=1.2B, # Trained Tokens=11T2025.11 | 20.7 | |
| LFM2-700M# Total Params=0.70B, # Trained Tokens=11T2025.11 | 20.56 | |
| Qwen3-0.6B# Total Params=0.6B, # Trained Tokens=36T2025.11 | 19.75 | |
| Gemma-3-1B# Total Params=1B, # Trained Tokens=2T2025.11 | 17.72 | |
| Llama-3.2-1B# Total Params=1.2B, # Trained Tokens=9T2025.11 | 16.86 | |
| LFM2-350M# Total Params=0.35B, # Trained Tokens=11T2025.11 | 16.41 |