Mathematical Reasoning on CMIMC
41.72AccuracyJustRL-Nemotron
Evaluation Results
| Method | Links | |
|---|---|---|
| JustRL-NemotronBackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@32, Dynamic Sampling=false, Training Steps=3440, Train Batch Size=256, Rollout N=8, Max Context Length=16k, Token Budget=1.1x10^8k2025.12 | 41.72 | |
| QuestABackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@32, Dynamic Sampling=true, Training Steps=2000, Train Batch Size=128, Rollout N=16, Max Context Length=32k, Token Budget=2.6x10^8k2025.12 | 41.48 | |
| BackboneBackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@322025.12 | 30.08 | |
| ProRL-V2Sampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 25.86 | |
| JustRL-DeepSeekSampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 25.63 | |
| DeepScaleR-1.5BSampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 21 | |
| Backbone (DeepSeek-R1-Distill-Qwen-1.5B)Sampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 12.89 |