Mathematical Reasoning on GSM8K Complex Samples (test)
26.1Avg@8 ScoreGLM-4.6
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GLM-4.6Model Category=Frontier model accessing by API2026.03 | 26.1 | 54.5 | 0.1 | |
| CLAUDE-4.5Model Category=Frontier model accessing by API2026.03 | 23.38 | 68.54 | 0.33 | |
| LLAMA3.2-3B-INSTModel Category=OSS foundation model2026.03 | 23.23 | 8.87 | 0.58 | |
| RETOOL-7BModel Category=RLVR-Trained model in TIR2026.03 | 23.04 | 28.02 | 1.25 | |
| LLAMA3.1-7B-INSTModel Category=OSS foundation model2026.03 | 21.87 | 17.63 | 0.99 | |
| SIMPLETIR-7BModel Category=RLVR-Trained model in TIR2026.03 | 19.26 | 52.74 | 1.83 | |
| ZEROTIR-7BModel Category=RLVR-Trained model in TIR2026.03 | 18.07 | 23.36 | 0.29 | |
| DEEPSEEK-R1Model Category=Frontier model accessing by API2026.03 | 16.92 | 52.84 | 0.69 | |
| QWEN2.5-7B-INSTModel Category=OSS foundation model2026.03 | 13.38 | 27.5 | 0.98 | |
| RETOOL-32BModel Category=RLVR-Trained model in TIR2026.03 | 12.99 | 12.24 | 2.9 | |
| GEMINI-3Model Category=Frontier model accessing by API2026.03 | 11.91 | 62.86 | 0.88 | |
| GPT-5.2Model Category=Frontier model accessing by API2026.03 | 11.22 | 17.03 | 0.2 | |
| QWEN3-32BModel Category=OSS foundation model2026.03 | 10.42 | 16.67 | 0.55 | |
| DEEPSEEK-V3.2Model Category=Frontier model accessing by API2026.03 | 9.01 | 20.06 | 1.26 | |
| QWEN3-8BModel Category=OSS foundation model2026.03 | 8.06 | 10.28 | 0.92 | |
| QWEN2.5-32B-INSTModel Category=OSS foundation model2026.03 | 7.5 | 12.5 | 1.58 |