Agentic Reasoning on HLE
41.6Overall ScoreChatGPT-Agent
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ChatGPT-Agent2026.02 | 41.6 | — | |
| InternAgent-1.5Base Model=Gemini-3-pro+o4-mini2026.02 | 40 | 40.87 | |
| InternAgent-1.5Base Model=o4-mini2026.02 | 34.52 | 36.1 | |
| Kimi-Researcher2026.02 | 26.9 | — | |
| Gemini DR2026.02 | 26.9 | — | |
| OpenAI DR2026.02 | 26.6 | — | |
| InternAgent-1.5Base Model=Qwen3-235B2026.02 | 14.84 | 15.04 | |
| MiroThinkerBase Model=MiroThinker-v1.5-30B2026.02 | — | 31 | |
| MiroThinkerBase Model=MiroThinker-v1.5-235B2026.02 | — | 39.2 | |
| Tongyi-DRBase Model=Tongyi-DR-30B2026.02 | — | 32.9 |