Mathematical Reasoning on Olympiad Bench (Pass@1)
62.1Pass@1EVOTD
Evaluation Results
| Method | Links | |
|---|---|---|
| EVOTDEvaluation Protocol=Pass@8, Training Iteration=Iter32026.05 | 62.1 | |
| EVOTDEvaluation Protocol=Pass@8, Training Iteration=Iter22026.05 | 61.5 | |
| EVOTDEvaluation Protocol=Pass@8, Training Iteration=Iter12026.05 | 59.7 | |
| Evol-InstructEvaluation Protocol=Pass@82026.05 | 59.1 | |
| Qwen3-4B (Thinking)Evaluation Protocol=Pass@82026.05 | 59 | |
| EVOTDEvaluation Protocol=Pass@1, Training Iteration=Iter32026.05 | 49.4 | |
| EVOTDEvaluation Protocol=Pass@1, Training Iteration=Iter22026.05 | 49.1 | |
| EVOTDEvaluation Protocol=Pass@1, Training Iteration=Iter12026.05 | 47.3 | |
| Evol-InstructEvaluation Protocol=Pass@12026.05 | 46.9 | |
| Qwen3-4B (Thinking)Evaluation Protocol=Pass@12026.05 | 45.6 | |
| BaselineBase Model=Qwen2.5-Math-7B-Instruct, Evaluation Protocol=Zero-shot2025.10 | 41.5 | |
| SFTBase Model=Qwen2.5-Math-7B-Instruct, Evaluation Protocol=Zero-shot2025.10 | 40.7 | |
| LightRBase Model=Qwen2.5-Math-7B-Instruct, Evaluation Protocol=Zero-shot2025.10 | 39 | |
| LightRBase Model=Qwen2.5-Math-1.5B-Instruct, Evaluation Protocol=Zero-shot2025.10 | 37.8 | |
| BaselineBase Model=Qwen2.5-Math-1.5B-Instruct, Evaluation Protocol=Zero-shot2025.10 | 37.5 | |
| SFTBase Model=Qwen2.5-Math-1.5B-Instruct, Evaluation Protocol=Zero-shot2025.10 | 37.5 | |
| LightRBase Model=DeepSeek-R1-Distill-Qwen-1.5B, Evaluation Protocol=Zero-shot2025.10 | 36.5 | |
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=false2025.09 | 36.1 | |
| Qwen2.5-7B + PPOBase Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=false2025.09 | 36.1 | |
| Qwen2.5-7B + PPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=true2025.09 | 36.1 | |
| Qwen2.5-7B + GRPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=true2025.09 | 35.4 | |
| SFTBase Model=Qwen2.5-Math-1.5B, Evaluation Protocol=Zero-shot2025.10 | 27.6 | |
| LightRBase Model=Qwen2.5-Math-1.5B, Evaluation Protocol=Zero-shot2025.10 | 27.1 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, RL Framework=None, VERL Integration=false2025.09 | 25.8 | |
| BaselineBase Model=Qwen2.5-Math-1.5B, Evaluation Protocol=Zero-shot2025.10 | 23.7 | |
| SFTBase Model=DeepSeek-R1-Distill-Qwen-1.5B, Evaluation Protocol=Zero-shot2025.10 | 21.2 | |
| SFTBase Model=Qwen2.5-Math-7B, Evaluation Protocol=Zero-shot2025.10 | 20.5 | |
| BaselineBase Model=DeepSeek-R1-Distill-Qwen-1.5B, Evaluation Protocol=Zero-shot2025.10 | 19.1 | |
| Llama-3.2-3B-Instruct + PPOBase Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=false2025.09 | 17.8 | |
| Llama-3.2-3B-Instruct + GRPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=true2025.09 | 17.6 | |
| Llama-3.2-3B-Instruct + PPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=true2025.09 | 17.3 | |
| LightRBase Model=Qwen2.5-Math-7B, Evaluation Protocol=Zero-shot2025.10 | 16.9 | |
| Llama-3.2-3B-Instruct + GRPOBase Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=false2025.09 | 16.7 | |
| BaselineBase Model=Qwen2.5-Math-7B, Evaluation Protocol=Zero-shot2025.10 | 16 | |
| Llama-3.2-3B-InstructBase Model=Llama-3.2-3B-Instruct, RL Framework=None, VERL Integration=false2025.09 | 12.7 |