Mathematical Reasoning on AIME24, AMC23, and MATH500
15AIME24 ScoreOAR-P
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OAR-PBackbone=Qwen2.5-Math-1.5B, Algorithm=GRPO, Method Variation=Outcome-grounded Advantage Reshaping (P)2026.01 | 15 | 58.7 | 75.5 | 49.7 | |
| OAR-GBackbone=Qwen2.5-Math-1.5B, Algorithm=GRPO, Method Variation=Outcome-grounded Advantage Reshaping (G)2026.01 | 14.4 | 58 | 75.1 | 49.2 | |
| Entropy AdvBackbone=Qwen2.5-Math-1.5B, Algorithm=GRPO, Credit Assignment=Entropy Advantage2026.01 | 14.2 | 58.2 | 74.5 | 48.9 | |
| GRPOBackbone=Qwen2.5-Math-1.5B, Algorithm=GRPO2026.01 | 13.8 | 56.7 | 73 | 47.8 | |
| Random AdvBackbone=Qwen2.5-Math-1.5B, Algorithm=GRPO, Credit Assignment=Random Advantage2026.01 | 13.8 | 57 | 72.6 | 47.8 | |
| Qwen2.5-Math-1.5BBackbone=Qwen2.5-Math-1.5B, Training=Pre-trained2026.01 | 2.9 | 20.2 | 40.3 | 21.1 |