Mathematical Reasoning on AIME 2025 (Accuracy@32, Average)
51.77Accuracy (Avg@32)SRaR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SRaRModel Size=32B2026.05 | 51.77 | — | |
| RaRModel Size=32B2026.05 | 50.94 | — | |
| DS-R1-Distill-Qwen-7B + CE-GPPOBackbone=Qwen-7B, Zero-shot mode=true, Time/Step (s)=292.52026.03 | 50.3 | 67.3 | |
| DS-R1-Distill-Qwen-7B + IC-DAPOBackbone=Qwen-7B, Zero-shot mode=true, Time/Step (s)=315.62026.03 | 49.8 | 67.8 | |
| SRaRModel Size=8B2026.05 | 49.06 | — | |
| DAPOModel Size=32B2026.05 | 47.81 | — | |
| DS-R1-Distill-Qwen-7B + DAPOBackbone=Qwen-7B, Zero-shot mode=true, Time/Step (s)=303.12026.03 | 45.9 | 65.3 | |
| RaRModel Size=8B2026.05 | 43.23 | — | |
| DAPOModel Size=8B2026.05 | 41.35 | — | |
| DS-R1-Distill-Qwen-7B + GRPOBackbone=Qwen-7B, Zero-shot mode=true, Time/Step (s)=305.62026.03 | 40.3 | 61.4 | |
| DS-R1-Distill-Qwen-7BBackbone=Qwen-7B, Zero-shot mode=true2026.03 | 39.1 | 61.8 | |
| RGR-GRPOModel Size=32B2026.05 | 36.46 | — | |
| RGR-GRPOModel Size=8B2026.05 | 34.69 | — | |
| DS-R1-Distill-Qwen-1.5B + IC-DAPOBackbone=Qwen-1.5B, Zero-shot mode=true, Time/Step (s)=477.22026.03 | 34.2 | 56.4 | |
| DS-R1-Distill-Qwen-1.5B + GSPOBackbone=Qwen-1.5B, Zero-shot mode=true, Time/Step (s)=437.32026.03 | 33.6 | 55.7 | |
| DS-R1-Distill-Qwen-1.5B + CE-GPPOBackbone=Qwen-1.5B, Zero-shot mode=true, Time/Step (s)=464.02026.03 | 32.5 | 55.7 | |
| GRPO-VPSModel Size=8B2026.05 | 32.4 | — | |
| GRPOModel Size=32B2026.05 | 30.21 | — | |
| DS-R1-Distill-Qwen-1.5B + DAPOBackbone=Qwen-1.5B, Zero-shot mode=true, Time/Step (s)=459.62026.03 | 28.4 | 53.9 | |
| DS-R1-Distill-Qwen-1.5B + GRPOBackbone=Qwen-1.5B, Zero-shot mode=true, Time/Step (s)=457.42026.03 | 28.1 | 50.3 | |
| GRPOModel Size=8B2026.05 | 27.71 | — | |
| GRPO-VPSModel Size=32B2026.05 | 26.56 | — | |
| DS-R1-Distill-Qwen-1.5B + CISPOBackbone=Qwen-1.5B, Zero-shot mode=true, Time/Step (s)=466.32026.03 | 25.1 | 48.8 | |
| DS-R1-Distill-Qwen-1.5BBackbone=Qwen-1.5B, Zero-shot mode=true2026.03 | 24.1 | 46.3 | |
| BaselineModel Size=8B2026.05 | 23.33 | — | |
| BaselineModel Size=32B2026.05 | 21.67 | — | |
| VIMPOReward condition=Clean, Training steps=Full training2026.06 | 20.8 | — | |
| GRPOReward condition=Clean, Training steps=Full training2026.06 | 17.6 | — | |
| GRPOModel=Qwen2.5-Math-7B, G=162026.06 | 15.94 | — | |
| VIMPOReward condition=25% reward flipping, Training steps=200 steps2026.06 | 14.9 | — | |
| Naive GRPOReward condition=Clean, Training steps=Full training2026.06 | 14.6 | — | |
| GRPOModel=Qwen2.5-Math-7B, G=22026.06 | 14.58 | — | |
| C-RF w/ NTFModel=Qwen2.5-Math-7B, G=12026.06 | 12.19 | — | |
| GRPOReward condition=25% reward flipping, Training steps=200 steps2026.06 | 11.7 | — | |
| GRPOModel=Qwen2.5-Math-1.5B, G=162026.06 | 10 | — | |
| C-RF w/ NTFModel=Qwen2.5-Math-1.5B, G=12026.06 | 9.38 | — | |
| GRPOModel=Qwen2.5-Math-1.5B, G=22026.06 | 8.13 | — | |
| RF++Model=Qwen2.5-Math-7B, G=82026.06 | 6.88 | — | |
| RF++Model=Qwen2.5-Math-1.5B, G=82026.06 | 6.46 | — | |
| RF++ w/ NTFModel=Qwen2.5-Math-7B, G=12026.06 | 6.15 | — | |
| Qwen3-4B-BaseReward condition=Clean, Training steps=Full training2026.06 | 3.6 | — | |
| RF++ w/ NTFModel=Qwen2.5-Math-1.5B, G=12026.06 | 2.6 | — |