Mathematical Reasoning on AIME25 (Pass@1)
72.5Pass@1RePO
Evaluation Results
| Method | Links | |
|---|---|---|
| RePOtraining_data=SuperGPQA subset (11k)2026.02 | 72.5 | |
| Qwen3-8Btraining_data=SuperGPQA subset (11k)2026.02 | 66.4 | |
| GRPOtraining_data=SuperGPQA subset (11k)2026.02 | 65.8 | |
| LUFFYtraining_data=SuperGPQA subset (11k)2026.02 | 64.1 | |
| Qwen3-4B + SCRLMax Generation Length=15,3602026.03 | 60.2 | |
| Qwen3-4B + TTRLMax Generation Length=15,3602026.03 | 59.3 | |
| Qwen3-4B + SCRLMax Generation Length=10,2402026.03 | 53.6 | |
| Qwen3-4BMax Generation Length=15,3602026.03 | 51.9 | |
| Qwen3-4B + TTRLMax Generation Length=10,2402026.03 | 50.9 | |
| Qwen3-1.7BModel Name=Qwen3-1.7B, Model Size=1.7B, Method Type=Baseline2026.04 | 47.5 | |
| Baseline ModelRollout Length=16K tokens2026.05 | 47.1 | |
| PieceHint-Nemotron-1.5BModel Name=Nemotron-1.5B, Model Size=1.5B, Method Type=PieceHint2026.04 | 43.7 | |
| DeepSeek-R1-Distill-32BModel Name=DeepSeek-R1-Distill-32B, Model Size=32B, Method Type=Baseline2026.04 | 39.5 | |
| Qwen3-4BMax Generation Length=10,2402026.03 | 38.9 | |
| PieceHint-Qwen3-1.7BModel Name=Qwen3-1.7B, Model Size=1.7B, Method Type=PieceHint2026.04 | 38.8 | |
| Qwen3-4BModel Name=Qwen3-4B, Model Size=4B, Method Type=Baseline2026.04 | 37.3 | |
| Nemotron-1.5BModel Name=Nemotron-1.5B, Model Size=1.5B, Method Type=Baseline2026.04 | 34.2 | |
| DeepSeek-R1-Distill-7BModel Name=DeepSeek-R1-Distill-7B, Model Size=7B, Method Type=Baseline2026.04 | 29.9 | |
| DAPO + PIPOModel=Qwen3-4B-Base2026.04 | 29.6 | |
| GRPO + PIPOModel=Qwen3-8B-Base2026.04 | 29.6 | |
| GSPO + PIPOModel=Qwen3-8B-Base2026.04 | 29.6 | |
| DAPO + PIPOModel=Qwen3-8B-Base2026.04 | 29.6 | |
| GSPO + PIPOModel=Qwen3-4B-Base2026.04 | 25.9 | |
| GRPOModel=Qwen3-8B-Base2026.04 | 25.9 | |
| DAPOModel=Qwen3-8B-Base2026.04 | 25.9 | |
| PieceHint-DeepSeek-1.5BModel Name=DeepSeek-R1-Distill-1.5B, Model Size=1.5B, Method Type=PieceHint2026.04 | 25.5 | |
| FP8 Rollout P3ORollout Length=16K tokens, Training Iteration=30, Rollout Precision=FP8, Training Precision=BF162026.05 | 25.4 | |
| DAPOModel=Qwen3-4B-Base2026.04 | 25 | |
| FP8 Rollout GRPORollout Length=16K tokens, Training Iteration=15, Rollout Precision=FP8, Training Precision=BF162026.05 | 25 | |
| FP8 Rollout P3ORollout Length=16K tokens, Training Iteration=15, Rollout Precision=FP8, Training Precision=BF162026.05 | 23.7 | |
| GRPO + PIPOModel=Qwen3-4B-Base2026.04 | 22.2 | |
| GSPOModel=Qwen3-8B-Base2026.04 | 22.2 | |
| DeepSeek-R1-Distill-1.5BModel Name=DeepSeek-R1-Distill-1.5B, Model Size=1.5B, Method Type=Baseline2026.04 | 21.8 | |
| GRPOModel=Qwen3-4B-Base2026.04 | 18.5 | |
| GSPOModel=Qwen3-4B-Base2026.04 | 18.5 | |
| P3ORollout Length=4K tokens, Clipping Parameter (epsilon)=∈ {0.2, 0.4, 0.6}2026.05 | 18.3 | |
| GRPO (clip avg)Rollout Length=4K tokens, Clipping Parameter (epsilon)=∈ {0.2, 0.4, 0.6}2026.05 | 16 | |
| SetPO+GRPOBase Model=Qwen2.5-Math-7B, Optimization Algorithm=GRPO2026.02 | 13.6 | |
| SetPO+DAPOBase Model=Qwen2.5-Math-7B, Optimization Algorithm=DAPO2026.02 | 13.5 | |
| Base ModelModel=Qwen3-8B-Base2026.04 | 11.1 | |
| DAPOBase Model=Qwen2.5-Math-7B2026.02 | 11 | |
| SetPO+GSPOBase Model=Qwen2.5-Math-7B, Optimization Algorithm=GSPO2026.02 | 10.1 | |
| GRPOBase Model=Qwen2.5-Math-7B2026.02 | 9.7 | |
| GSPOBase Model=Qwen2.5-Math-7B2026.02 | 7.9 | |
| Base ModelModel=Qwen3-4B-Base2026.04 | 7.4 | |
| Qwen2.5-Math-7BModel Type=Base Model2026.02 | 4 | |
| Baseline ModelRollout Length=4K tokens2026.05 | 3.3 | |
| FP8 Rollout GRPORollout Length=16K tokens, Training Iteration=30, Rollout Precision=FP8, Training Precision=BF162026.05 | 0 |