Mathematical Reasoning on Minerva (Accuracy)
51.47AccuracyJustRL-DeepSeek
Evaluation Results
| Method | Links | |
|---|---|---|
| JustRL-DeepSeekSampling strategy=@4, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 51.47 | |
| BroRLSampling strategy=@4, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 49.08 | |
| ProRL-V2Sampling strategy=@4, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 49.03 | |
| AZR-R-TAPBase Model Type=Base2026.03 | 45.7 | |
| AZR-R-TAPBase Model Type=Coder2026.03 | 44.3 | |
| R1-Distill-Qwen-7B-R-TAP2026.03 | 42.3 | |
| DeepScaleR-1.5BSampling strategy=@4, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 39.34 | |
| Oat-Zero-7B-R-TAP2026.03 | 37.2 | |
| PRIME-Zero-7B2026.03 | 36 | |
| R1-Distill-Qwen-7B @ 8kTraining iterations=8k2026.03 | 35.9 | |
| Backbone (DeepSeek-R1-Distill-Qwen-1.5B)Sampling strategy=@4, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 34.65 | |
| R1-Distill-Qwen-1.5B-R-TAP @ 8kTraining iterations=8k2026.03 | 31.9 | |
| OpenReasoner-Zero-7B @ 3kTraining iterations=3k2026.03 | 31.6 | |
| OpenReasoner-Zero-7B @ 8kTraining iterations=8k2026.03 | 31.6 | |
| Oat-Zero-1.5B-R-TAP2026.03 | 31.2 | |
| Oat-Zero-7B2026.03 | 30.1 | |
| Qwen2.5-Math-7B-Instruct2026.03 | 29.8 | |
| SimpleRL-Zero-7B2026.03 | 27.6 | |
| Qwen2.5-Math-1.5B-Instruct2026.03 | 26.5 | |
| Oat-Zero-1.5B2026.03 | 25.7 | |
| R1-Distill-Qwen-1.5B @ 8kTraining iterations=8k2026.03 | 25 | |
| R1-Distill-Qwen-7B @ 3kTraining iterations=3k2026.03 | 23 | |
| Qwen2.5-Math-7B2026.03 | 21.3 | |
| R1-Distill-Qwen-1.5B @ 3kTraining iterations=3k2026.03 | 16.3 | |
| Qwen2.5-Math-1.5B2026.03 | 15.1 |