Mathematical Reasoning on SAT-Math (accuracy)
97.66SAT Math AccuracyL1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| L1Base Model=DeepSeek-1.5B2026.02 | 97.66 | 2,377 | |
| GPT-42024.05 | 96.9 | — | |
| AbstRaLModel=Qwen2.5-Math-7B-Instruct, GSM Train Method=AbstRaL, Evaluation Protocol=Zero-shot2025.06 | 93.8 | — | |
| CRTBase Model=DeepSeek-1.5B2026.02 | 93.16 | 1,582 | |
| O1-PrunerBase Model=DeepSeek-1.5B, Ratio=4.02026.02 | 91.21 | 1,500 | |
| ACPOBase Model=DeepSeek-1.5B2026.02 | 91.02 | 1,282 | |
| Ori-SFTModel=Qwen2.5-Math-7B-Instruct, GSM Train Method=Ori-SFT, Evaluation Protocol=Zero-shot2025.06 | 90.6 | — | |
| CoAModel=Qwen2.5-Math-7B-Instruct, GSM Train Method=CoA, Evaluation Protocol=Zero-shot2025.06 | 90 | — | |
| DeepSeek-1.5BModel Type=Base Model2026.02 | 89.84 | 1,373 | |
| CoT-RLModel=Qwen2.5-Math-7B-Instruct, GSM Train Method=CoT-RL, Evaluation Protocol=Zero-shot2025.06 | 89.8 | — | |
| Qwen-1.5Parameters=110B2024.05 | 87.5 | — | |
| Qwen-1.5Parameters=72B2024.05 | 87.5 | — | |
| MAmmoTH2Parameters=8B, Variant=Plus2024.05 | 87.5 | — | |
| O1-PrunerBase Model=DeepSeek-1.5B, Ratio=1.02026.02 | 86.91 | 1,225 | |
| ShorterBetterBase Model=DeepSeek-1.5B2026.02 | 84.77 | 200 | |
| DeepSeekMathParameters=7B2024.05 | 84.4 | — | |
| DeepSeekMathParameters=7B, Variant=Instruct2024.05 | 84.4 | — | |
| MAmmoTH2Parameters=7B, Variant=Plus2024.05 | 84.4 | — | |
| JiuZhang3.0Parameters=8B2024.05 | 84.4 | — | |
| ThinkPruneBase Model=DeepSeek-1.5B, Constraint=5002026.02 | 81.45 | 609 | |
| ThinkPruneBase Model=DeepSeek-1.5B, Constraint=1k2026.02 | 81.45 | 759 | |
| MAmmoTH2Parameters=8x7B, Variant=Plus2024.05 | 81.2 | — | |
| JiuZhang3.0Parameters=8x7B2024.05 | 81.2 | — | |
| JiuZhang3.0Parameters=7B2024.05 | 81.2 | — | |
| ChatGPT2024.05 | 78.1 | — | |
| Rho-1-MathParameters=7B2024.05 | 75 | — | |
| LlemmaParameters=34B2024.05 | 71.9 | — | |
| GemmaParameters=7B2024.05 | 71.9 | — | |
| MixtralParameters=8x7B2024.05 | 65.6 | — | |
| Intern-MathParameters=20B2024.05 | 65.6 | — | |
| LlemmaParameters=7B2024.05 | 62.5 | — | |
| AbstRaLModel=Qwen2.5-0.5B-Instruct, GSM Train Method=AbstRaL, Evaluation Protocol=Zero-shot2025.06 | 62.5 | — | |
| MistralParameters=7B2024.05 | 56.2 | — | |
| LLAMA-3Parameters=8B2024.05 | 56.2 | — | |
| Ori-SFTModel=Qwen2.5-0.5B-Instruct, GSM Train Method=Ori-SFT, Evaluation Protocol=Zero-shot2025.06 | 46.9 | — | |
| CoT-RLModel=Qwen2.5-0.5B-Instruct, GSM Train Method=CoT-RL, Evaluation Protocol=Zero-shot2025.06 | 43.8 | — | |
| CoAModel=Qwen2.5-0.5B-Instruct, GSM Train Method=CoA, Evaluation Protocol=Zero-shot2025.06 | 43.8 | — | |
| Llama-3-SynEevaluation_mode=Few-shot2024.07 | 43.64 | — | |
| DCLM-7Bevaluation_mode=Few-shot2024.07 | 41.36 | — | |
| MAmmoTH2-8Bevaluation_mode=Few-shot2024.07 | 41.36 | — | |
| Mistral-7B-v0.3evaluation_mode=Few-shot2024.07 | 40.45 | — | |
| Llama-3-8Bevaluation_mode=Few-shot2024.07 | 38.64 | — | |
| Llama-3-Chinese-8Bevaluation_mode=Few-shot2024.07 | 36.82 | — | |
| CENTRIFUGEFoundation Model=TinyLlama-1.1B, Training Dataset=OWM (Filter 50%)2025.02 | 25 | — | |
| Galactica-6.7Bevaluation_mode=Few-shot2024.07 | 23.18 | — | |
| No FinetuningFoundation Model=TinyLlama-1.1B, Training Dataset=NA2025.02 | 21.9 | — | |
| Regular FinetuningFoundation Model=TinyLlama-1.1B, Training Dataset=OWM (Full)2025.02 | 18.8 | — |