Mathematical Reasoning on GSM8K Simple Samples (test)
99.29Accuracy (avg@8)DEEPSEEK-V3.2
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DEEPSEEK-V3.2Model Category=Frontier model accessing by API2026.03 | 99.29 | 94.18 | 0.98 | |
| GPT-5.2Model Category=Frontier model accessing by API2026.03 | 99.28 | 95.01 | 0.1 | |
| QWEN3-8BModel Category=OSS foundation model2026.03 | 99.22 | 97.6 | 0.56 | |
| QWEN3-32BModel Category=OSS foundation model2026.03 | 99.21 | 97.6 | 0.37 | |
| QWEN2.5-32B-INSTModel Category=OSS foundation model2026.03 | 98.52 | 97.07 | 1.47 | |
| RETOOL-32BModel Category=RLVR-Trained model in TIR2026.03 | 98.08 | 91.79 | 3.09 | |
| GEMINI-3Model Category=Frontier model accessing by API2026.03 | 97.97 | 86.48 | 0.86 | |
| QWEN2.5-7B-INSTModel Category=OSS foundation model2026.03 | 96.99 | 91.66 | 1.03 | |
| DEEPSEEK-R1Model Category=Frontier model accessing by API2026.03 | 96.44 | 79.66 | 0.68 | |
| GLM-4.6Model Category=Frontier model accessing by API2026.03 | 95.22 | 90.24 | 0.09 | |
| ZEROTIR-7BModel Category=RLVR-Trained model in TIR2026.03 | 92.62 | 83.75 | 0.21 | |
| RETOOL-7BModel Category=RLVR-Trained model in TIR2026.03 | 92.27 | 87.21 | 1.19 | |
| LLAMA3.1-7B-INSTModel Category=OSS foundation model2026.03 | 91.95 | 64.24 | 1 | |
| SIMPLETIR-7BModel Category=RLVR-Trained model in TIR2026.03 | 89.86 | 96.93 | 1.2 | |
| LLAMA3.2-3B-INSTModel Category=OSS foundation model2026.03 | 88.58 | 39.43 | 0.7 | |
| CLAUDE-4.5Model Category=Frontier model accessing by API2026.03 | 70.98 | 72.84 | 0.44 |