Mathematical Reasoning on MATH, AIME25, AMC, MINERVA, KAOYAN, OLYMPIAD, and CN_MATH24
92.2MATH AccuracyOpenThinker2-32B
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| OpenThinker2-32BStudent Model=Qwen2.5-32B-Instruct, Training Data=1.04M samples (DeepSeek-R1 teacher)2025.10 | 92.2 | 56.7 | 90 | 32.4 | 64.8 | 64 | 83.3 | 69.1 | |
| Local HighestStudent Model=Qwen2.5-32B-Instruct, Selection Strategy=Local average log probabilities2025.10 | 90.2 | 66.7 | 100 | 35.3 | 65.3 | 67.3 | 83.3 | 72.6 | |
| LIMO-32BStudent Model=Qwen2.5-32B-Instruct, Training Data=LIMO prompt set (DeepSeek-R1 teacher)2025.10 | 89.6 | 43.3 | 92.5 | 34.6 | 61.8 | 63 | 80 | 66.4 | |
| Global HighestStudent Model=Qwen2.5-32B-Instruct, Selection Strategy=Global average log probabilities2025.10 | 87.6 | 43.3 | 82.5 | 33.1 | 59.2 | 63.6 | 73.3 | 63.2 | |
| Sky-T1-32B-PreviewStudent Model=Qwen2.5-32B-Instruct, Training Data=17K example dataset (QWQ-32B model)2025.10 | 87.6 | 20 | 75 | 30.1 | 55.8 | 50.7 | 53.3 | 53.2 | |
| Original ModelStudent Model=Qwen2.5-32B-Instruct2025.10 | 82.4 | 13.3 | 70 | 29.8 | 42.2 | 47.1 | 23.3 | 44.5 |