Mathematical Reasoning on AIME Complex Samples 24/25 (AVG@8, Tool Usage)
23.86AVG@8GEMINI-3
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GEMINI-3Model Category=Frontier model2026.03 | 23.86 | 39.2 | 3.76 | |
| GLM-4.6Model Category=Frontier model2026.03 | 12.78 | 16.39 | 0.11 | |
| QWEN3-8BModel Category=OSS foundation models2026.03 | 10.19 | 17.13 | 1.8 | |
| CLAUDE-4.5Model Category=Frontier model2026.03 | 9.8 | 16.89 | 0.8 | |
| DEEPSEEK-R1Model Category=Frontier model2026.03 | 8.55 | 12.83 | 0.25 | |
| QWEN2.5-7B-INSTModel Category=OSS foundation models2026.03 | 5.39 | 4.17 | 0.42 | |
| DEEPSEEK-V3.2Model Category=Frontier model2026.03 | 5 | 51.25 | 5.5 | |
| RETOOL-32BModel Category=RLVR-Trained models in TIR2026.03 | 4.22 | 23.78 | 1.5 | |
| QWEN2.5-32B-INSTModel Category=OSS foundation models2026.03 | 3.32 | 10.2 | 0.73 | |
| RETOOL-7BModel Category=RLVR-Trained models in TIR2026.03 | 2.84 | 17.65 | 1.63 | |
| SIMPLETIR-7B*Model Category=RLVR-Trained models in TIR, Suppresses invalid tool-use=true2026.03 | 2.5 | 26.5 | 5.18 | |
| ZEROTIR-7BModel Category=RLVR-Trained models in TIR2026.03 | 2.45 | 6.25 | 1.15 | |
| LLAMA3.2-3B-INSTModel Category=OSS foundation models2026.03 | 1.75 | 0.66 | 1.17 | |
| LLAMA3.1-7B-INSTModel Category=OSS foundation models2026.03 | 0.91 | 3.18 | 0.95 | |
| GPT-5.2Model Category=Frontier model2026.03 | 0 | 0 | 0.47 |