Mathematical Reasoning on AIME Simple Samples 24/25 (AVG@8, Tool Usage)
99.75AVG@8 Success RateGPT-5.2
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5.2Model Category=Frontier model2026.03 | 99.75 | 52.25 | 1.33 | |
| DEEPSEEK-V3.2Model Category=Frontier model2026.03 | 95.25 | 73.75 | 3.12 | |
| DEEPSEEK-R1Model Category=Frontier model2026.03 | 90.38 | 66.35 | 0.13 | |
| CLAUDE-4.5Model Category=Frontier model2026.03 | 88.64 | 70.45 | 0.76 | |
| GEMINI-3Model Category=Frontier model2026.03 | 88.16 | 71.05 | 1.92 | |
| GLM-4.6Model Category=Frontier model2026.03 | 85 | 69.17 | 0.07 | |
| RETOOL-32BModel Category=RLVR-Trained models in TIR2026.03 | 84.45 | 80.21 | 2.15 | |
| LLAMA3.2-3B-INSTModel Category=OSS foundation models2026.03 | 83.33 | 20.83 | 1.92 | |
| LLAMA3.1-7B-INSTModel Category=OSS foundation models2026.03 | 80 | 57.5 | 1 | |
| QWEN3-8BModel Category=OSS foundation models2026.03 | 74.24 | 75 | 1.83 | |
| QWEN2.5-7B-INSTModel Category=OSS foundation models2026.03 | 73.61 | 68.06 | 0.46 | |
| QWEN2.5-32B-INSTModel Category=OSS foundation models2026.03 | 72.73 | 72.73 | 1.39 | |
| SIMPLETIR-7B*Model Category=RLVR-Trained models in TIR, Suppresses invalid tool-use=true2026.03 | 72.5 | 87.5 | 2.38 | |
| ZEROTIR-7BModel Category=RLVR-Trained models in TIR2026.03 | 62.5 | 75 | 1.03 | |
| RETOOL-7BModel Category=RLVR-Trained models in TIR2026.03 | 56.93 | 87.5 | 1.12 |