Mathematical Reasoning on AIME 24 (Pass@1 and Pass@8)
0.233Pass@1GRPO (Ts : 1.5)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPO (Ts : 1.5)Sampling temperature (Ts)=1.5, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 0.233 | 0.367 | |
| TAMPOMaximum response length=6k tokens, Training algorithm=TAMPO2026.02 | 0.233 | 0.4 | |
| GRPO (Ts : 0.9)Sampling temperature (Ts)=0.9, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 0.2 | 0.3 | |
| GRPO (Ts : 1.2)Sampling temperature (Ts)=1.2, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 0.2 | 0.333 | |
| GRPO (Ts : 0.9 -> 1.5)Sampling temperature (Ts)=0.9 -> 1.5, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 0.167 | 0.3 | |
| DS-Qwen-1.5BMaximum response length=6k tokens2026.02 | 0.133 | 0.267 |