Mathematical Reasoning on AMC 2023 (Pass@1, Pass@32, Avg)
73.6Pass@1Qwen3-8B-Base
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-8B-BaseTraining=RLAD (ours), Context=30K2026.02 | 73.6 | 98 | 66.5 | |
| Qwen3-8B-BaseTraining=GRPO, Context=30K2026.02 | 72.9 | 94.8 | 61 | |
| Qwen3-8B-BaseTraining=KDRL, Context=30K2026.02 | 72.8 | 95.8 | 64.9 | |
| ISPOBase Model=Qwen2.5-Math-7B, Decoding Strategy=Greedy (T=0)2026.06 | 72.5 | — | — | |
| Scaf-GRPOBase Model=Qwen2.5-Math-7B, Decoding Strategy=Greedy (T=0)2026.06 | 70 | — | — | |
| ISPOBase Model=Qwen2.5-7B, Decoding Strategy=Greedy (T=0)2026.06 | 68.3 | — | — | |
| Qwen3-8B-BaseTraining=SFT, Context=30K2026.02 | 67.9 | 92.8 | 56 | |
| ProGRPOBase Model=Qwen2.5-7B, Decoding Strategy=Greedy (T=0)2026.06 | 67.2 | — | — | |
| Vanilla GRPOBase Model=Qwen2.5-7B, Decoding Strategy=Greedy (T=0)2026.06 | 65.5 | — | — | |
| ISPOBase Model=Qwen2.5-Math-1.5B, Decoding Strategy=Greedy (T=0)2026.06 | 62.5 | — | — | |
| Scaf-GRPOBase Model=Qwen2.5-Math-1.5B, Decoding Strategy=Greedy (T=0)2026.06 | 60 | — | — | |
| Vanilla GRPOBase Model=Qwen2.5-Math-7B, Decoding Strategy=Greedy (T=0)2026.06 | 60 | — | — | |
| Scaf-GRPOBase Model=Qwen2.5-7B, Decoding Strategy=Greedy (T=0)2026.06 | 60 | — | — | |
| Dr. GRPOBase Model=Qwen2.5-Math-7B, Decoding Strategy=Greedy (T=0)2026.06 | 56.6 | — | — | |
| PACRBase Model=Qwen2.5-Math-7B, Decoding Strategy=Greedy (T=0)2026.06 | 56.1 | — | — | |
| PACRBase Model=Qwen2.5-Math-1.5B, Decoding Strategy=Greedy (T=0)2026.06 | 49.4 | — | — | |
| Vanilla GRPOBase Model=Qwen2.5-Math-1.5B, Decoding Strategy=Greedy (T=0)2026.06 | 47.5 | — | — | |
| Dr. GRPOBase Model=Qwen2.5-Math-1.5B, Decoding Strategy=Greedy (T=0)2026.06 | 47 | — | — | |
| Qwen3-1.7B-BaseTraining=RLAD (ours), Context=30K2026.02 | 42 | 83.7 | 39 | |
| Qwen3-1.7B-BaseTraining=GRPO, Context=30K2026.02 | 41.2 | 83.8 | 36.5 | |
| Qwen3-1.7B-BaseTraining=KDRL, Context=30K2026.02 | 41.1 | 83.6 | 37.2 | |
| Base modelBase Model=Qwen2.5-Math-7B, Decoding Strategy=Greedy (T=0)2026.06 | 38.6 | — | — | |
| Base modelBase Model=Qwen2.5-Math-1.5B, Decoding Strategy=Greedy (T=0)2026.06 | 32.5 | — | — | |
| Base modelBase Model=Qwen2.5-7B, Decoding Strategy=Greedy (T=0)2026.06 | 32.4 | — | — | |
| Qwen3-8B-BaseTraining=-, Context=30K2026.02 | 30.5 | 88.6 | 36.8 | |
| Qwen3-1.7B-BaseTraining=-, Context=30K2026.02 | 16.3 | 70.8 | 22.2 |