Mathematical Reasoning on AIME25 (pass@1 (%))
66.67Pass@1PREPO
Evaluation Results
| Method | Links | |
|---|---|---|
| PREPOBase Model=Qwen3-4B, Training Strategy=PREPO, # Rollouts=348K2025.11 | 66.67 | |
| Random SelectionBase Model=Qwen3-4B, Training Strategy=Random Selection, # Rollouts=553K2025.11 | 60 | |
| GRESOBase Model=Qwen3-4B, Training Strategy=GRESO, # Rollouts=472K2025.11 | 56.67 | |
| Qwen3-4BBase Model=Qwen3-4B, Training Strategy=Base, # Rollouts=–2025.11 | 30 | |
| Random SelectionBase Model=Qwen2.5-Math-1.5B, Training Strategy=Random Selection, # Rollouts=3.0M2025.11 | 20 | |
| PREPOBase Model=Qwen2.5-Math-1.5B, Training Strategy=PREPO, # Rollouts=1.1M2025.11 | 20 | |
| GRESOBase Model=Qwen2.5-Math-7B, Training Strategy=GRESO, # Rollouts=654K2025.11 | 18.33 | |
| GRESOBase Model=Qwen2.5-Math-1.5B, Training Strategy=GRESO, # Rollouts=2.5M2025.11 | 15.38 | |
| PREPOBase Model=Qwen2.5-Math-7B, Training Strategy=PREPO, # Rollouts=540K2025.11 | 12.81 | |
| PREPOBase Model=Qwen2.5-7B, Training Strategy=PREPO, # Rollouts=304K2025.11 | 10.21 | |
| Random SelectionBase Model=Qwen2.5-Math-7B, Training Strategy=Random Selection, # Rollouts=905K2025.11 | 10 | |
| GRESOBase Model=Qwen2.5-7B, Training Strategy=GRESO, # Rollouts=680K2025.11 | 9.22 | |
| Qwen2.5-Math-7BBase Model=Qwen2.5-Math-7B, Training Strategy=Base, # Rollouts=–2025.11 | 9.17 | |
| Random SelectionBase Model=Qwen2.5-7B, Training Strategy=Random Selection, # Rollouts=716K2025.11 | 6.98 | |
| Qwen2.5-Math-1.5BBase Model=Qwen2.5-Math-1.5B, Training Strategy=Base, # Rollouts=–2025.11 | 3.54 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, Training Strategy=Base, # Rollouts=–2025.11 | 1.25 |