Mathematical Reasoning on AIME 24 (Avg@32, #Token)
66.8Average Score (Top-32)DC-GRPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DC-GRPOBackbone=Qwen3-30B-A3B, Precision=BF162026.06 | 66.8 | — | 71.8 | |
| DC-GRPOBackbone=Qwen3-8B, Precision=BF162026.06 | 63.2 | — | 71 | |
| GRPOBackbone=Qwen3-30B-A3B, Precision=BF162026.06 | 60.2 | — | 69 | |
| GRPOBackbone=Qwen3-8B, Precision=BF162026.06 | 58.8 | — | 67.3 | |
| GR³Category=Performance-oriented RL2026.03 | 45.2 | 8,381 | — | |
| SFT→RLBase Model=Qwen2.5-Math-7B, Source=verl, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 40.4 | — | — | |
| GRPOCategory=Performance-oriented RL2026.03 | 39.6 | 13,054 | — | |
| SRFT [13]Base Model=Qwen2.5-Math-7B, Source=reported, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 35.3 | — | — | |
| DLER-R1-1.5BCategory=Length-oriented RL2026.03 | 34.3 | 3,839 | — | |
| AdaptThink-1.5BCategory=Length-oriented RL2026.03 | 34.2 | 9,204 | — | |
| SFTBase Model=Qwen2.5-Math-7B, Source=verl, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 33.6 | — | — | |
| HPT [15]Base Model=Qwen2.5-Math-7B, Source=reported, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 33 | — | — | |
| Prefix-RFT [14]Base Model=Qwen2.5-Math-7B, Source=reported, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 31.8 | — | — | |
| Laser-DE-L4096-1.5BCategory=Length-oriented RL2026.03 | 30.1 | 5,770 | — | |
| DeepSeek-R1-Distill-1.5BCategory=Initial model2026.03 | 30 | 16,531 | — | |
| LUFFY [11]Base Model=Qwen2.5-Math-7B, Source=reported, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 29.4 | — | — | |
| ReLIFT [12]Base Model=Qwen2.5-Math-7B, Source=reported, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 28.3 | — | — | |
| ReLIFT [12]Base Model=Qwen2.5-Math-7B, Source=reproduction attempt, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 25.6 | — | — | |
| LUFFY [11]Base Model=Qwen2.5-Math-7B, Source=reproduction attempt, Sampling Temperature=0.6, Max Response Length=8192 tokens, Verifier=Math-Verify2026.04 | 25 | — | — | |
| LCR1-1.5BCategory=Length-oriented RL2026.03 | 23.5 | 9,071 | — |