Expert Question Answering on HLE
11.2AccuracyQwQ-32B
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| QwQ-32BModel Type=Natively-trained reasoning LM, Parameters=32B2026.06 | 11.2 | 54.5 | 82 | 60.7 | 70 | 71 | 74.2 | 66 | |
| DeepSeek-R1-8BModel Type=Distilled LRM, Parameters=8B2026.06 | 6.3 | 68 | 71.4 | 72.6 | 65.3 | 76 | 78.5 | 65.1 |