Mathematical Reasoning on OlympiadBench (Pass@1, Pass@8)
61.2Pass@8GRPO (Ts : 1.5)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPO (Ts : 1.5)Sampling temperature (Ts)=1.5, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 61.2 | 39 | |
| TAMPOMaximum response length=6k tokens, Training algorithm=TAMPO2026.02 | 60.7 | 39.6 | |
| GRPO (Ts : 1.2)Sampling temperature (Ts)=1.2, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 60.5 | 38.1 | |
| GRPO (Ts : 0.9 -> 1.5)Sampling temperature (Ts)=0.9 -> 1.5, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 60.4 | 41 | |
| GRPO (Ts : 0.9)Sampling temperature (Ts)=0.9, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 59.4 | 38.7 | |
| DS-Qwen-1.5BMaximum response length=6k tokens2026.02 | 59 | 38.4 | |
| Qwen2.5-7B-InstructTraining Pipeline=PSFT → GRPO2025.08 | 58.5 | — | |
| Qwen2.5-7B-InstructTraining Pipeline=SFT → GRPO2025.08 | 58.09 | — | |
| SFT-KLBackbone=Qwen2.5-7B-Instruct2025.08 | 52.7 | — | |
| SFTBackbone=Qwen2.5-7B-Instruct2025.08 | 52.35 | — | |
| Qwen2.5-7B-InstructTraining Pipeline=SFT2025.08 | 52.35 | — | |
| PSFTwarm-upBackbone=Qwen2.5-7B-Instruct2025.08 | 52.3 | — | |
| PSFTBackbone=Qwen2.5-7B-Instruct2025.08 | 51.5 | — | |
| Qwen2.5-7B-InstructTraining Pipeline=PSFT2025.08 | 51.5 | — | |
| Llama3.1-8B-InstructTraining Pipeline=PSFT → GRPO2025.08 | 47.74 | — | |
| Llama3.1-8B-InstructTraining Pipeline=SFT → GRPO2025.08 | 47.13 | — | |
| PSFTwarm-upBackbone=Llama3.1-8B-Instruct2025.08 | 42.07 | — | |
| SFTBackbone=Llama3.1-8B-Instruct2025.08 | 41.43 | — | |
| Llama3.1-8B-InstructTraining Pipeline=SFT2025.08 | 41.43 | — | |
| BaseBackbone=Qwen2.5-7B-Instruct2025.08 | 39.72 | — | |
| PSFTBackbone=Llama3.1-8B-Instruct2025.08 | 39.61 | — | |
| Llama3.1-8B-InstructTraining Pipeline=PSFT2025.08 | 39.61 | — | |
| SFT-KLBackbone=Llama3.1-8B-Instruct2025.08 | 39.19 | — | |
| BaseBackbone=Llama3.1-8B-Instruct2025.08 | 17.09 | — | |
| Curriculum SFTModels=Qwen2.5-14B-Instruct2026.03 | — | 57.53 | |
| Curriculum SFTModels=Qwen3-4B-Base2026.03 | — | 39.92 | |
| HEALModels=Qwen2.5-14B-Instruct2026.03 | — | 62.28 | |
| HEALModels=Qwen3-4B-Base2026.03 | — | 42.17 | |
| LIMOModels=Qwen2.5-14B-Instruct2026.03 | — | 35.14 | |
| LIMOModels=Qwen3-4B-Base2026.03 | — | 38.43 | |
| OriginModels=Qwen2.5-14B-Instruct2026.03 | — | 42.45 | |
| OriginModels=Qwen3-4B-Base2026.03 | — | 38.16 | |
| SFTModels=Qwen2.5-14B-Instruct2026.03 | — | 54.12 | |
| SFTModels=Qwen3-4B-Base2026.03 | — | 39.83 |