Mathematical Reasoning on Olympiad Bench (Accuracy)
76.71AccuracySRaR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SRaRModel Size=32B2026.05 | 76.71 | — | |
| SkyWork-OR1-Math-7BTeacher Model=SkyWork-OR1-Math-7B2026.03 | 73.5 | — | |
| SRaRModel Size=8B2026.05 | 73.15 | — | |
| SkyWork-OR1-7BTeacher Model=SkyWork-OR1-7B2026.03 | 72.5 | — | |
| RGR-GRPOModel Size=32B2026.05 | 72.4 | — | |
| DAPOModel Size=32B2026.05 | 72.11 | — | |
| RaRModel Size=32B2026.05 | 69.29 | — | |
| DAPOModel Size=8B2026.05 | 67.8 | — | |
| RaRModel Size=8B2026.05 | 67.51 | — | |
| RGR-GRPOModel Size=8B2026.05 | 66.91 | — | |
| GRPOModel Size=32B2026.05 | 66.02 | — | |
| PODSModel=Qwen3-4B, Selection Ratio=50%2026.05 | 66 | — | |
| Full DataModel=Qwen3-4B, Selection Ratio=100%2026.05 | 65.7 | — | |
| PODSModel=Qwen3-8B, Selection Ratio=50%2026.05 | 65.1 | — | |
| GRPO-VPSModel Size=32B2026.05 | 64.54 | — | |
| Full DataModel=Qwen3-8B, Selection Ratio=100%2026.05 | 64.5 | — | |
| GRPOModel Size=8B2026.05 | 63.2 | — | |
| GRPO-VPSModel Size=8B2026.05 | 60.24 | — | |
| GSPO + REINFORCE++Objective=GSPO, G (Rollouts per prompt)=1, Steps=300, Time=13.0h2026.05 | 59.9 | — | |
| SLATbackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 59.6 | 3,714 | |
| Originalbackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 59.1 | 8,093 | |
| AdaptThinkbackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 59.1 | 5,945 | |
| GSPO (G=8)Objective=GSPO, G (Rollouts per prompt)=8, Steps=300, Time=14.4h2026.05 | 58.8 | — | |
| BaselineModel Size=32B2026.05 | 58.75 | — | |
| GSPO VanillaObjective=GSPO, G (Rollouts per prompt)=1, Steps=300, Time=13.0h2026.05 | 58.6 | — | |
| DASTbackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 58.3 | 7,390 | |
| TLMREbackbone=DeepSeek-R1-Distill-Qwen-7B, alpha=0.12026.05 | 58 | 6,289 | |
| GPG + BASISObjective=GPG, G (Rollouts per prompt)=1, Steps=150, Time=6.4h2026.05 | 57.4 | — | |
| REOPOLDTeacher Model=SkyWork-OR1-Math-7B2026.03 | 57.3 | — | |
| LC-R1backbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 57.3 | 4,156 | |
| TLMREbackbone=DeepSeek-R1-Distill-Qwen-7B, alpha=0.22026.05 | 57.2 | 5,679 | |
| GPG (G=8)Objective=GPG, G (Rollouts per prompt)=8, Steps=300, Time=14.0h2026.05 | 57 | — | |
| GSPO + BASISObjective=GSPO, G (Rollouts per prompt)=1, Steps=150, Time=7.2h2026.05 | 56.5 | — | |
| Legislator-Executor (Ours)Model=Qwen3-14B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 56.3 | — | |
| BaselineModel Size=8B2026.05 | 56.23 | — | |
| RKLTeacher Model=SkyWork-OR1-Math-7B2026.03 | 56 | — | |
| GRPO + BASISObjective=GRPO, G (Rollouts per prompt)=1, Steps=150, Time=8.3h2026.05 | 55.9 | — | |
| GPG + REINFORCE++Objective=GPG, G (Rollouts per prompt)=1, Steps=300, Time=12.9h2026.05 | 55.8 | — | |
| L1-Maxbackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 55.8 | 2,854 | |
| SFTTeacher Model=SkyWork-OR1-Math-7B2026.03 | 55.6 | — | |
| Qwen3-14B + GRPOBackbone=Qwen3-14B, Training=GRPO2026.05 | 55.2 | 4,824 | |
| Qwen3-14B + SLATBackbone=Qwen3-14B, Training=SLAT2026.05 | 54.8 | 2,106 | |
| ConSteer-RLBackbone=Qwen3-8B-Base2026.06 | 53.9 | — | |
| GRPO (G=8)Objective=GRPO, G (Rollouts per prompt)=8, Steps=300, Time=15.5h2026.05 | 53.7 | — | |
| REOPOLDTeacher Model=SkyWork-OR1-7B2026.03 | 53.3 | — | |
| RREDCoT-Nanocheckpoints=2, context size=10242026.06 | 53.1 | — | |
| S1KModel=Qwen3-14B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 52.6 | — | |
| Open-RScheckpoints=3, context size=10242026.06 | 52.4 | — | |
| Qwen2.5-1.5B R1-Distilled (base)context size=10242026.06 | 52.1 | — | |
| R1-Qwen-1.5B + EvoCoTBackbone=R1-Qwen-1.5B, Algorithm=EvoCoT2025.08 | 52 | — | |
| GRPOBackbone=Qwen3-8B-Base2026.06 | 51.7 | — | |
| R1-Qwen-1.5B + DeepScaleR(GRPO)Backbone=R1-Qwen-1.5B, Algorithm=DeepScaleR(GRPO)2025.08 | 51.6 | — | |
| LIMOModel=Qwen3-14B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 50.8 | — | |
| Legislator-Executor (Ours)Model=Qwen2.5-7B-QwQ, Decoding Strategy=Zero-shot greedy2026.04 | 50.7 | — | |
| Legislator-Executor (Ours)Model=Qwen3-8B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 50.5 | — | |
| GRPOTeacher Model=SkyWork-OR1-Math-7B2026.03 | 49.8 | — | |
| GRPO + REINFORCE++Objective=GRPO, G (Rollouts per prompt)=1, Steps=300, Time=17.7h2026.05 | 48.7 | — | |
| RKLTeacher Model=SkyWork-OR1-7B2026.03 | 47.9 | — | |
| R1-Qwen-1.5B + SFTBackbone=R1-Qwen-1.5B, Algorithm=SFT2025.08 | 47.4 | — | |
| Legislator-Executor (Ours)Model=Qwen2.5-Math-7B, Decoding Strategy=Zero-shot greedy2026.04 | 47 | — | |
| ConSteer-RLBackbone=Qwen3-4B-Base2026.06 | 47 | — | |
| GRPOBackbone=Qwen3-4B-Base2026.06 | 46.7 | — | |
| BaseModel=Qwen3-14B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 46.2 | — | |
| Qwen2.5-7B + Open-ReasonerBackbone=Qwen2.5-7B, Algorithm=Open-Reasoner2025.08 | 45.6 | — | |
| LIMOModel=Qwen3-8B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 44.9 | — | |
| S1KModel=Qwen2.5-7B-QwQ, Decoding Strategy=Zero-shot greedy2026.04 | 44.7 | — | |
| Legislator-Executor (Ours)Model=Qwen3-4B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 44.6 | — | |
| S1KModel=Qwen3-8B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 44.4 | — | |
| LIMOModel=Qwen2.5-7B-QwQ, Decoding Strategy=Zero-shot greedy2026.04 | 44.1 | — | |
| R1-Distill-Qwen-1.5BStudent Model=DeepSeek-R1-Distill-Qwen-1.5B2026.03 | 43.6 | — | |
| R1-Qwen-1.5BBackbone=R1-Qwen-1.5B2025.08 | 43.3 | — | |
| BaseModel=Qwen3-8B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 41.8 | — | |
| Qwen2.5-7B + PRIME (380K)Backbone=Qwen2.5-7B, Algorithm=PRIME, Training Data Size=380K2025.08 | 41.8 | — | |
| Uni-DPOModel=Qwen2.5-Math 7B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 41.5 | — | |
| Token-ALPevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 40.77 | — | |
| S1KModel=Qwen3-4B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 40.6 | — | |
| BaselineBackbone=Qwen3-8B-Base2026.06 | 40.6 | — | |
| Seq-ALPevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 40.28 | — | |
| GRPO-SGBase Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 40.12 | — | |
| Token-MISevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 40.06 | — | |
| LIMOModel=Qwen3-4B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 40 | — | |
| Qwen3-I(4B)Model Family=Qwen3, Target Model (TL)=ℐ (4B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 39.8 | — | |
| BaseModel=Qwen2.5-7B-QwQ, Decoding Strategy=Zero-shot greedy2026.04 | 39.6 | — | |
| Qwen2.5-Math-7B w/ CADFTBackbone=Qwen2.5-Math-7B, Fine-tuning Protocol=CADFT, Evaluation Protocol=Average@162026.04 | 39.5 | — | |
| SimPOModel=Qwen2.5-Math 7B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 39.3 | — | |
| Seq-MISevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 39.06 | — | |
| GRPOevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 38.82 | — | |
| Seq-Bypassevaluation_protocol=average@32, temperature=1.0, max_tokens=4096, setting=Single-turn2026.03 | 38.52 | — | |
| Legislator-Executor (Ours)Model=Qwen2.5-7B, Decoding Strategy=Zero-shot greedy2026.04 | 38.5 | — | |
| ConSteer-RLBackbone=Qwen2.5-Math-7B2026.06 | 38.5 | — | |
| GRPO-SGBase Model=Qwen2.5-7B, Training Recipe=DSR (DeepScaleR)2025.10 | 38.48 | — | |
| GRPOBase Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 38.45 | — | |
| 80/20Base Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 38.2 | — | |
| Qwen2.5-7B + SimpleRL(GRPO)Backbone=Qwen2.5-7B, Algorithm=SimpleRL(GRPO)2025.08 | 38.1 | — | |
| ARBase Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 37.95 | — | |
| BaseModel=Qwen3-4B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 37.9 | — | |
| LIMOModel=Qwen2.5-Math-7B, Decoding Strategy=Zero-shot greedy2026.04 | 37.9 | — | |
| DPOModel=Qwen2.5-Math 7B, zero-shot chain-of-thought prompting=true, greedy decoding=true2025.06 | 37.9 | — | |
| GRPOBackbone=Qwen2.5-Math-7B2026.06 | 37.9 | — | |
| Qwen3-I(14B)Model Family=Qwen3, Target Model (TL)=ℐ (14B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 37.8 | — |