Mathematical Reasoning on AIME24, AIME25, MATH, MINERVA, GPQA, and GSM8K (test)
68.75AIME24 ScoreFully Trained Model
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Fully Trained ModelBase Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 68.75 | 62 | 97.5 | 61.74 | 54.16 | 98.25 | 73.65 | |
| AlphaRLStage=50%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 68.25 | 62.5 | 97.25 | 62.58 | 54.16 | 97.75 | 73.75 | |
| AlphaRLStage=40%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 67.5 | 61.5 | 97.25 | 62.35 | 53.67 | 97.75 | 73.34 | |
| TrainingStage=50%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 65.75 | 58.85 | 96.5 | 60.65 | 52.27 | 97.25 | 71.87 | |
| AlphaRLStage=30%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 64.5 | 57.25 | 96.17 | 60.07 | 51.33 | 96.5 | 70.97 | |
| TrainingStage=40%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 63.25 | 55.25 | 95 | 58.45 | 49.86 | 96.5 | 69.74 | |
| AlphaRLStage=20%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 60 | 52.25 | 95.5 | 57.85 | 48.25 | 95.5 | 68.23 | |
| TrainingStage=30%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 59.5 | 51.5 | 95.2 | 56.22 | 47.25 | 95.5 | 67.52 | |
| AlphaRLStage=10%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 56.5 | 44.5 | 94.5 | 54.25 | 43.85 | 94.75 | 64.73 | |
| TrainingStage=20%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 56.5 | 45.75 | 93.75 | 53.75 | 44.25 | 94.75 | 64.79 | |
| TrainingStage=10%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 52.5 | 41.75 | 92.75 | 51.5 | 41.17 | 94.25 | 62.32 | |
| AlphaRLStage=5%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 51.92 | 40.08 | 92.75 | 51.35 | 40.75 | 93 | 61.64 | |
| TrainingStage=5%, Base Model=Deepseek-R1-Distill-Qwen-7B (After Distillation)2025.10 | 50.83 | 39.17 | 91.25 | 50 | 38.89 | 93 | 60.69 | |
| AlphaRLStage=40%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 40.75 | 30 | 87.5 | 41.5 | 44.25 | 94.5 | 56.75 | |
| AlphaRLStage=50%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 39.12 | 30.75 | 87.75 | 42.25 | 43.75 | 94.25 | 56.98 | |
| Fully Trained ModelBase Model=Llama3-8B-Instruct (After SFT)2025.10 | 38.75 | 31.25 | 87.5 | 42.75 | 43.75 | 95 | 56.83 | |
| TrainingStage=50%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 31.11 | 27.11 | 86.25 | 37.5 | 42.25 | 92.5 | 52.79 | |
| AlphaRLStage=30%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 30.64 | 27.02 | 85.75 | 38.85 | 42.13 | 93.5 | 53.98 | |
| TrainingStage=40%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 28.33 | 23.67 | 84.25 | 37.5 | 40.75 | 92.25 | 51.46 | |
| TrainingStage=30%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 23.67 | 20.33 | 81.5 | 35.25 | 38.25 | 91.5 | 48.75 | |
| AlphaRLStage=20%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 22.67 | 20 | 84.75 | 34.75 | 37.25 | 92.25 | 48.94 | |
| TrainingStage=20%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 20 | 15 | 78 | 31.07 | 34.75 | 89.5 | 45.06 | |
| AlphaRLStage=10%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 15.97 | 13.06 | 77.75 | 31.76 | 34.54 | 89.25 | 43.72 | |
| TrainingStage=10%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 12.63 | 9.16 | 73 | 28.37 | 31.75 | 86.25 | 40.19 | |
| AlphaRLStage=5%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 12.36 | 9.02 | 71.25 | 28.37 | 31.25 | 86.5 | 39.79 | |
| TrainingStage=5%, Base Model=Llama3-8B-Instruct (After SFT)2025.10 | 9.58 | 6.67 | 68.75 | 26.44 | 29.75 | 85.25 | 37.74 |