Open-ended Writing on ArenaHard
50AccuracyOfficial Instruct Model
Evaluation Results
| Method | Links | |
|---|---|---|
| Official Instruct ModelEvaluation Category=Official Instruct Model2026.05 | 50 | |
| Opus-4.7 (xHigh)Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 33.84 | |
| GLM-4.7 & ANDESEvaluation Category=Proposed Method Comparison, Base Model=Qwen3-1.7B2026.05 | 12.94 | |
| GLM-4.7 (Scaffold-only)Evaluation Category=Proposed Method Comparison, Base Model=Qwen3-1.7B2026.05 | 5.7 | |
| Opus-4.6 (1M)Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 3.21 | |
| Gemini-3.1-ProEvaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 2.27 | |
| MiniMax-M2.5Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 1.76 | |
| GPT-5.2Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 1.27 | |
| Opus-4.6Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 1.15 | |
| Base Model (Qwen3-1.7B)Evaluation Category=Zero-Shot2026.05 | 0.91 | |
| MiniMax-M2.1Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0.91 | |
| Qwen3-MaxEvaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0.91 | |
| Kimi-K2-ThinkingEvaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0.91 | |
| GPT-5.4 (High)Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0.91 | |
| GLM-4.7 (OpenCode)Evaluation Category=Proposed Method Comparison, Base Model=Qwen3-1.7B2026.05 | 0.91 | |
| Sonnet-4.5Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0.21 | |
| GPT-5.1-Codex-MaxEvaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0.14 |