Instruction Following on IFEval (Strict Accuracy)
90Strict AccuracyLlama-3.3-70B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-3.3-70B-InstructMaximum output tokens=32,7682025.07 | 90 | |
| SMCSMaximum output tokens=32,7682025.07 | 90 | |
| Claude-3.7-Sonnet(2025-02-19)Maximum output tokens=32,7682025.07 | 88 | |
| Qwen-2.5-72B-InstructMaximum output tokens=32,7682025.07 | 86.3 | |
| GPT-4.1(2025-04-14)Maximum output tokens=32,7682025.07 | 86 | |
| GLM-Z1-32B-0414Maximum output tokens=32,7682025.07 | 84.3 | |
| Qwen3-32BMaximum output tokens=32,7682025.07 | 83 | |
| Llama-3.3-Nemotron-Super-49B-v1Maximum output tokens=32,7682025.07 | 82.7 | |
| GPT-4o(2024-08-06)Maximum output tokens=32,7682025.07 | 82.3 | |
| QwQ-32BMaximum output tokens=32,7682025.07 | 82.3 | |
| GPT-o3-mini(2025-01-31)Maximum output tokens=32,7682025.07 | 82 | |
| TeleChat2-35B-32KMaximum output tokens=32,7682025.07 | 82 | |
| Gemma-3-27b-itMaximum output tokens=32,7682025.07 | 81 | |
| DeepSeek-R1-Distill-Llama-70BMaximum output tokens=32,7682025.07 | 80.7 | |
| Claude-3.5-Sonnet(2024-06-20)Maximum output tokens=32,7682025.07 | 80.3 | |
| Qwen2.5-Coder-32B-InstructMaximum output tokens=32,7682025.07 | 80.3 | |
| Qwen2.5-32b-InstructMaximum output tokens=32,7682025.07 | 78.7 | |
| EXAONE-Deep-32BMaximum output tokens=32,7682025.07 | 76.3 | |
| DeepSeek-R1-Distill-Qwen-32BMaximum output tokens=32,7682025.07 | 74.3 | |
| HuatuoGPT-o1-72BMaximum output tokens=32,7682025.07 | 74 | |
| InternLM2.5-20B-ChatMaximum output tokens=32,7682025.07 | 64.7 |