Complex Reasoning on Humanity's Last Exam (HLE)
18.4Pass@1 Scoregemini-2.5-pro
Evaluation Results
| Method | Links | |
|---|---|---|
| gemini-2.5-pro# Params=–, Scale Category=Large Scale (> 70B)2026.03 | 18.4 | |
| Qwen3-235B-A22B-Thinking-2507# Params=235B, Scale Category=Large Scale (> 70B)2026.03 | 18.2 | |
| o4-mini (high)# Params=–, Scale Category=Large Scale (> 70B)2026.03 | 18.1 | |
| DeepSeek-R1-0528# Params=671B, Scale Category=Large Scale (> 70B)2026.03 | 17.7 | |
| Qwen3-235B-A22B# Params=235B, Scale Category=Large Scale (> 70B)2026.03 | 11.8 | |
| o3-mini (medium)# Params=–, Scale Category=Large Scale (> 70B)2026.03 | 10.3 | |
| Qwen3-4B-Thinking-2507 + CHIMERA# Params=4B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 9 | |
| Qwen3-32B# Params=32B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 8.9 | |
| DeepSeek-R1# Params=671B, Scale Category=Large Scale (> 70B)2026.03 | 8.5 | |
| Qwen3-4B-Thinking-2507# Params=4B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 7.3 | |
| DeepSeek-R1-0528-Qwen3-8B# Params=8B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 6.9 | |
| DeepSeek-R1-Distill-Llama-70B# Params=70B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 5.2 | |
| Qwen3-4B-Thinking-2507 + OpenScience# Params=4B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 4.6 |