Mathematical Reasoning on AIME25, AMC23, MATH500, Minerva Aggregate
72.16Average ScoreGRPO w/ Structure Reward
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPO w/ Structure RewardRL Category=Label-free RL, Algorithm=GRPO, Reward Method=Structure Reward2026.03 | 72.16 | 7.65 | |
| GRPO w/ Ground-TruthRL Category=Ground-Truth RL (Oracle), Algorithm=GRPO, Reward Method=Ground-Truth2026.03 | 71.66 | 7.15 | |
| GRPO w/ Entropy Minimization (EMPO)RL Category=Label-free RL, Algorithm=GRPO, Reward Method=Entropy Minimization (EMPO)2026.03 | 71.45 | 6.94 | |
| GRPO w/ Majority Voting (TTRL)RL Category=Label-free RL, Algorithm=GRPO, Reward Method=Majority Voting (TTRL)2026.03 | 71.12 | 6.61 | |
| PPO w/ Structure RewardRL Category=Label-free RL, Algorithm=PPO, Reward Method=Structure Reward2026.03 | 70.38 | 5.87 | |
| PPO w/ Ground-TruthRL Category=Ground-Truth RL (Oracle), Algorithm=PPO, Reward Method=Ground-Truth2026.03 | 70.18 | 5.67 | |
| Base2026.03 | 64.51 | — |