Mathematical Reasoning on AIME 25 (Avg@32, Average)
34.38Avg@32BuPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BuPOBackbone=Qwen3-8B2025.12 | 34.38 | 66.36 | |
| GRPOBackbone=Qwen3-8B2025.12 | 33.54 | 64.23 | |
| RLOOBackbone=Qwen3-8B2025.12 | 33.02 | 63.36 | |
| BuPOBackbone=Qwen3-4B2025.12 | 31.15 | 58.51 | |
| Reinforce++Backbone=Qwen3-8B2025.12 | 31.15 | 60.41 | |
| GRPOBackbone=Qwen3-4B2025.12 | 28.85 | 55.08 | |
| PPOBackbone=Qwen3-4B2025.12 | 27.6 | 55.22 | |
| RLOOBackbone=Qwen3-4B2025.12 | 24.79 | 54 | |
| PPOBackbone=Qwen3-8B2025.12 | 22.6 | 58.41 | |
| VanillaBackbone=Qwen3-8B2025.12 | 19.17 | 48.49 | |
| Reinforce++Backbone=Qwen3-4B2025.12 | 18.65 | 45.03 | |
| VanillaBackbone=Qwen3-4B2025.12 | 18.6 | 47.44 | |
| BuPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 6.77 | 27.79 | |
| Reinforce++Backbone=Llama-OctoThinker-8B-Base2025.12 | 3.75 | 26.43 | |
| GRPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 2.19 | 24.11 | |
| RLOOBackbone=Llama-OctoThinker-8B-Base2025.12 | 1.56 | 22.18 | |
| PPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 1.04 | 22.82 | |
| BuPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 0.42 | 19.59 | |
| PPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 0.31 | 16.69 | |
| RLOOBackbone=Llama-OctoThinker-3B-Base2025.12 | 0.21 | 17.84 | |
| Reinforce++Backbone=Llama-OctoThinker-3B-Base2025.12 | 0.1 | 5.27 | |
| GRPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 0.1 | 18.58 | |
| VanillaBackbone=Llama-OctoThinker-8B-Base2025.12 | 0.1 | 3.75 | |
| VanillaBackbone=Llama-OctoThinker-3B-Base2025.12 | 0 | 1.68 |