Instruction Following on Instruction Following tasks
78.3ScoreGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Generation Mode=LLM-in-Sandbox2026.01 | 78.3 | 7 | |
| DeepSeek-V3.2-ThinkingGeneration Mode=LLM-in-Sandbox2026.01 | 74.7 | 14.4 | |
| MiniMax-M2Generation Mode=Standard LLM2026.01 | 73 | — | |
| Claude-Sonnet-4.5-ThinkGeneration Mode=LLM-in-Sandbox2026.01 | 72 | 12.7 | |
| GPT-5Generation Mode=Standard LLM2026.01 | 71.3 | — | |
| Kimi-K2-ThinkingGeneration Mode=LLM-in-Sandbox2026.01 | 68.7 | 3.7 | |
| Kimi-K2-ThinkingGeneration Mode=Standard LLM2026.01 | 65 | — | |
| MiniMax-M2Generation Mode=LLM-in-Sandbox2026.01 | 61.3 | -11.7 | |
| DeepSeek-V3.2-ThinkingGeneration Mode=Standard LLM2026.01 | 60.3 | — | |
| Claude-Sonnet-4.5-ThinkGeneration Mode=Standard LLM2026.01 | 59.3 | — | |
| Qwen3-Coder-30B-A3BGeneration Mode=LLM-in-Sandbox2026.01 | 40.3 | 5.6 | |
| Qwen3-Coder-30B-A3BGeneration Mode=Standard LLM2026.01 | 34.7 | — | |
| Qwen3-4B-Instruct-2507Generation Mode=Standard LLM2026.01 | 33.7 | — | |
| Qwen3-4B-Instruct-2507Generation Mode=LLM-in-Sandbox2026.01 | 29 | -4.7 |