Mathematical Reasoning on (A24, A25, AMC, MATH, Minerva)
52.29A24 ScoreSPPO
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SPPOBackbone=DeepSeek-R1-Distill-Qwen-7B, Critic Configuration=Small Critic (1.5B)2026.04 | 52.29 | 34.58 | 87.19 | 89.88 | 28.86 | 58.56 | |
| SPPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.04 | 50.83 | 35 | 86.25 | 90.13 | 28.35 | 58.11 | |
| ReMaxBackbone=DeepSeek-R1-Distill-Qwen-7B2026.04 | 49.38 | 31.25 | 86.56 | 90.28 | 27.99 | 57.09 | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-7B, Sample Count (N)=82026.04 | 47.08 | 35 | 86.25 | 90.15 | 28.74 | 57.44 | |
| RLOOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.04 | 46.67 | 32.5 | 86.88 | 90.35 | 28.72 | 57.02 | |
| PPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.04 | 45.2 | 35.42 | 85.31 | 88.48 | 27.8 | 56.44 | |
| Base ModelBackbone=DeepSeek-R1-Distill-Qwen-7B2026.04 | 41.25 | 26.67 | 79.38 | 87.2 | 27.94 | 52.49 | |
| SPPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 34.17 | 25.83 | 74.38 | 83.78 | 22.15 | 48.06 | |
| ReMaxBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 31.67 | 25.42 | 71.88 | 84.38 | 20.4 | 46.74 | |
| RLOOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 30.42 | 21.67 | 72.81 | 84.1 | 21.73 | 46.15 | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Sample Count (N)=82026.04 | 30 | 26.25 | 73.13 | 83.88 | 22.15 | 47.08 | |
| Base ModelBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 27.5 | 21.67 | 71.56 | 83.73 | 20.35 | 44.96 | |
| PPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 27.5 | 20.83 | 70.63 | 81.38 | 19.89 | 44.06 |