Mathematical Reasoning on BRUMO
80AccuracyGRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 80 | |
| SPSBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 80 | |
| QuestABackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@32, Dynamic Sampling=true, Training Steps=2000, Train Batch Size=128, Rollout N=16, Max Context Length=32k, Token Budget=2.6x10^8k2025.12 | 67.5 | |
| JustRL-NemotronBackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@32, Dynamic Sampling=false, Training Steps=3440, Train Batch Size=256, Rollout N=8, Max Context Length=16k, Token Budget=1.1x10^8k2025.12 | 66.88 | |
| GSPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 66.67 | |
| BackboneBackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@322025.12 | 61.67 | |
| DeepSeek-R1-Distill-Qwen-1.5BType=Base Model2026.04 | 60 | |
| DAPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 60 | |
| JustRL-DeepSeekSampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 52.71 | |
| ProRL-V2Sampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 47.29 | |
| DeepScaleR-1.5BSampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 40 | |
| Backbone (DeepSeek-R1-Distill-Qwen-1.5B)Sampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 30.94 |