Hard Reasoning and Language Evaluation on HLE
54AccuracyKimi-K2.6
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Kimi-K2.6Tools usage=with tools, Parameter Count=1T-A32B2026.06 | 54 | — | — | — | |
| GLM-5.1Tools usage=with tools, Parameter Count=744B-A40B2026.06 | 50.4 | — | — | — | |
| Qwen-3.5Tools usage=with tools, Parameter Count=397B-17B2026.06 | 48.3 | — | — | — | |
| DS-v4-ProTools usage=with tools, Parameter Count=1.6T-A49B2026.06 | 48.2 | — | — | — | |
| DS-v4-FlashTools usage=with tools, Parameter Count=284B-A13B2026.06 | 45.1 | — | — | — | |
| DS-v4-ProTools usage=no tools, Parameter Count=1.6T-A49B2026.06 | 37.7 | — | — | — | |
| Nemotron 3 UltraTools usage=with tools, Parameter Count=550B-A55B2026.06 | 37.4 | — | — | — | |
| WebClipper (Hybrid)Category=Pruning Method, Base Model=Tongyi-DeepResearch, Variant=Hybrid2026.02 | 36.1 | 0.495 | 21.07 | 13,532 | |
| Tongyi-DeepResearchCategory=Open-sourced Agent, Evaluation Environment=Unified2026.02 | 35.8 | 0.487 | 23.92 | 13,664 | |
| WebClipper (Eff)Category=Pruning Method, Base Model=Tongyi-DeepResearch, Variant=Efficiency2026.02 | 35.3 | 0.492 | 18.6 | 11,458 | |
| Prompt ControlCategory=Pruning Method, Base Model=Tongyi-DeepResearch2026.02 | 34.9 | 0.479 | 23.91 | 14,107 | |
| Kimi-K2.6Tools usage=no tools, Parameter Count=1T-A32B2026.06 | 34.8 | — | — | — | |
| Coarse PruneCategory=Pruning Method, Base Model=Tongyi-DeepResearch2026.02 | 32.7 | 0.467 | 18.03 | 11,851 | |
| DS-v4-FlashTools usage=no tools, Parameter Count=284B-A13B2026.06 | 32.2 | — | — | — | |
| Qwen-3.5Tools usage=no tools, Parameter Count=397B-17B2026.06 | 28.5 | — | — | — | |
| GLM-5.1Tools usage=no tools, Parameter Count=744B-A40B2026.06 | 27.2 | — | — | — | |
| Nemotron 3 UltraTools usage=no tools, Parameter Count=550B-A55B2026.06 | 26.7 | — | — | — | |
| OpenAI DeepResearchCategory=Close-sourced System2026.02 | 26.6 | — | — | — | |
| OpenAI o3Category=Close-sourced System2026.02 | 24.9 | — | — | — | |
| MiniMax-2.7Tools usage=no tools, Parameter Count=230B-A10B2026.06 | 23.1 | — | — | — | |
| Claude-4-SonnetCategory=Close-sourced System2026.02 | 20.3 | — | — | — | |
| Qwen3-235B-A22B-Instruct-2507Category=Open-sourced Agent, Evaluation Environment=Unified2026.02 | 19.9 | 0.327 | 7.45 | 2,960 | |
| Kimi-K2-Instruct-0905Category=Open-sourced Agent, Evaluation Environment=Unified2026.02 | 14.6 | 0.253 | 5.17 | 2,349 | |
| DeepSeek-R1-671BCategory=Open-sourced Agent, Evaluation Environment=Unified2026.02 | 13.7 | 0.239 | 5.89 | 2,394 | |
| WebExplorerCategory=Open-sourced Agent2026.02 | 11.6 | 0.203 | 15.52 | 6,579 |