Mathematical Reasoning on Minerva (avg@8)
62.45Avg@8GRPO w/ Ground-Truth
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPO w/ Ground-TruthRL Category=Ground-Truth RL (Oracle), Algorithm=GRPO, Reward Method=Ground-Truth2026.03 | 62.45 | |
| GRPO w/ Structure RewardRL Category=Label-free RL, Algorithm=GRPO, Reward Method=Structure Reward2026.03 | 61.99 | |
| GRPO w/ Majority Voting (TTRL)RL Category=Label-free RL, Algorithm=GRPO, Reward Method=Majority Voting (TTRL)2026.03 | 61.86 | |
| GRPO w/ Entropy Minimization (EMPO)RL Category=Label-free RL, Algorithm=GRPO, Reward Method=Entropy Minimization (EMPO)2026.03 | 61.44 | |
| PPO w/ Structure RewardRL Category=Label-free RL, Algorithm=PPO, Reward Method=Structure Reward2026.03 | 61.08 | |
| PPO w/ Ground-TruthRL Category=Ground-Truth RL (Oracle), Algorithm=PPO, Reward Method=Ground-Truth2026.03 | 59.74 | |
| Base2026.03 | 53.45 | |
| Qwen2.5-7B-InstructTraining Pipeline=PSFT → GRPO2025.08 | 46.83 | |
| Qwen2.5-7B-InstructTraining Pipeline=SFT → GRPO2025.08 | 45.27 | |
| Qwen2.5-7B-InstructTraining Pipeline=SFT2025.08 | 43.66 | |
| Qwen2.5-7B-InstructTraining Pipeline=PSFT2025.08 | 43.33 | |
| Llama3.1-8B-InstructTraining Pipeline=PSFT → GRPO2025.08 | 37.55 | |
| Llama3.1-8B-InstructTraining Pipeline=SFT → GRPO2025.08 | 33.09 | |
| Llama3.1-8B-InstructTraining Pipeline=PSFT2025.08 | 32.4 | |
| Llama3.1-8B-InstructTraining Pipeline=SFT2025.08 | 32.17 |