Mathematical Reasoning on HMMT (Solved Rate %)
80.01Solved RateGPT-OSS-120B
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-OSS-120BSampling Strategy=4-sample mean2026.04 | 80.01 | |
| Aryabhata 2Sampling Strategy=4-sample mean2026.04 | 78.96 | |
| GPT-OSS-20BSampling Strategy=4-sample mean2026.04 | 77.42 | |
| GPT-5 MiniSampling Strategy=4-sample mean2026.04 | 70.97 | |
| Nemotron 3 Nano 30B A3BSampling Strategy=4-sample mean2026.04 | 65.86 | |
| GPT-5 NanoSampling Strategy=4-sample mean2026.04 | 63.98 | |
| Gemini 2.5 FlashSampling Strategy=4-sample mean2026.04 | 59.13 | |
| Qwen3-30B-A3B (Thinking)Sampling Strategy=4-sample mean2026.04 | 51.88 | |
| ReflexionModel=Qwen3 32B think, Context Window=32k2025.05 | 0.6 | |
| ICRLModel=Qwen3 32B think, Context Window=32k2025.05 | 0.6 | |
| Self-RefineModel=Qwen3 32B think, Context Window=32k2025.05 | 0.5666 | |
| BaseModel=Qwen3 32B think, Context Window=32k2025.05 | 0.52 | |
| ICRLModel=Qwen3 32B, Context Window=32k2025.05 | 0.3333 | |
| ReflexionModel=Qwen3 32B, Context Window=32k2025.05 | 0.2333 | |
| ICRLModel=Llama 4 Maverick, Context Window=32k2025.05 | 0.2 | |
| Self-RefineModel=Qwen3 32B, Context Window=32k2025.05 | 0.1666 | |
| Self-RefineModel=Llama 4 Maverick, Context Window=32k2025.05 | 0.1333 | |
| Self-RefineModel=Phi-4, Context Window=16k2025.05 | 0.1333 | |
| ReflexionModel=Phi-4, Context Window=16k2025.05 | 0.1333 | |
| ICRLModel=Phi-4, Context Window=16k2025.05 | 0.1333 | |
| ReflexionModel=Llama 4 Maverick, Context Window=32k2025.05 | 0.1 | |
| BaseModel=Qwen3 32B, Context Window=32k2025.05 | 0.0914 | |
| BaseModel=Llama 4 Maverick, Context Window=32k2025.05 | 0.085 | |
| BaseModel=Phi-4, Context Window=16k2025.05 | 0.0555 |