Mathematical Reasoning on OlympiadBench (pass@1, pass@5)
62.8Pass@1SFT→RL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SFT→RLBase Model=Qwen2.5-Math-7B, Source=verl2026.04 | 62.8 | — | |
| SRFT [13]Base Model=Qwen2.5-Math-7B, Source=reported2026.04 | 58.3 | — | |
| ReLIFT [12]Base Model=Qwen2.5-Math-7B, Source=reported2026.04 | 57.3 | — | |
| LUFFY [11]Base Model=Qwen2.5-Math-7B, Source=reported2026.04 | 57.2 | — | |
| HPT [15]Base Model=Qwen2.5-Math-7B, Source=reported2026.04 | 56.9 | — | |
| SFTBase Model=Qwen2.5-Math-7B, Source=verl2026.04 | 55.7 | — | |
| Prefix-RFT [14]Base Model=Qwen2.5-Math-7B, Source=reported2026.04 | 55.7 | — | |
| ReLIFT [12]Base Model=Qwen2.5-Math-7B, Source=reproduction attempt2026.04 | 55.6 | — | |
| LUFFY [11]Base Model=Qwen2.5-Math-7B, Source=reproduction attempt2026.04 | 50.8 | — | |
| APMPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 46.6 | — | |
| GMPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 44.8 | — | |
| DAPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 43.8 | — | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 43.2 | — | |
| APMPOBackbone=Qwen2.5-Math-1.5B-Instruct2026.04 | 42.4 | — | |
| DAPOBackbone=Qwen2.5-Math-1.5B-Instruct2026.04 | 40.4 | — | |
| GRPOBackbone=Qwen2.5-Math-1.5B-Instruct2026.04 | 39 | — | |
| GMPOBackbone=Qwen2.5-Math-1.5B-Instruct2026.04 | 38.7 | — | |
| BaseBackbone=Qwen2.5-Math-1.5B-Instruct2026.04 | 35.2 | — | |
| INFOTREE2026.05 | 34.9 | — | |
| APMPOBackbone=Qwen2.5-3B-Instruct2026.04 | 33.2 | — | |
| DAPOBackbone=Qwen2.5-3B-Instruct2026.04 | 32.6 | — | |
| GMPOBackbone=Qwen2.5-3B-Instruct2026.04 | 32.2 | — | |
| Tree-GRPO2026.05 | 31.7 | — | |
| GRPOBackbone=Qwen2.5-3B-Instruct2026.04 | 31.5 | — | |
| BaseBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 30.9 | — | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.04 | 29.1 | — | |
| Flat GRPO2026.05 | 28.4 | — | |
| Base ModelBackbone=Llama-3.2-3B-Instruct2025.10 | 0.1132 | 0.1956 | |
| FABackbone=Llama-3.2-3B-Instruct, Synthetic Data Strategy=GUIDEDSAMPLING - Final Answer (FA)2025.10 | 0.1121 | 0.2021 | |
| CAABackbone=Llama-3.2-3B-Instruct, Synthetic Data Strategy=GUIDEDSAMPLING - Concepts and Answer (CAA)2025.10 | 0.1076 | 0.2047 | |
| ToTBackbone=Llama-3.2-3B-Instruct, Synthetic Data Strategy=Tree-of-Thought2025.10 | 0.0919 | 0.1836 | |
| RSBackbone=Llama-3.2-3B-Instruct, Synthetic Data Strategy=Repeated Sampling2025.10 | 0.0642 | 0.1083 | |
| STaRBackbone=Llama-3.2-3B-Instruct, Synthetic Data Strategy=Self-Taught Reasoner2025.10 | 0.0582 | 0.1062 |