Multi-disciplinary Reasoning on HLE
37.7AccuracyGemini-3.0
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-3.0variant=Pro, protocol=Pass@12025.12 | 37.7 | 15,000 | |
| DeepSeek-V3.2variant=Speciale, protocol=Pass@12025.12 | 30.6 | 35,000 | |
| GPT-5variant=High, protocol=Pass@12025.12 | 26.3 | 15,000 | |
| DeepSeek-V3.2variant=Thinking, protocol=Pass@12025.12 | 25.1 | 21,000 | |
| Kimi-K2variant=Thinking, protocol=Pass@12025.12 | 23.9 | 24,000 |