Mathematical Reasoning on AIME24 (Avg@32, Average)
54.06Avg@32BuPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BuPOBackbone=Qwen3-8B2025.12 | 54.06 | 66.36 | |
| GRPOBackbone=Qwen3-8B2025.12 | 49.48 | 64.23 | |
| RLOOBackbone=Qwen3-8B2025.12 | 46.67 | 63.36 | |
| Reinforce++Backbone=Qwen3-8B2025.12 | 41.77 | 60.41 | |
| PPOBackbone=Qwen3-8B2025.12 | 37.81 | 58.41 | |
| BuPOBackbone=Qwen3-4B2025.12 | 36.88 | 58.51 | |
| PPOBackbone=Qwen3-4B2025.12 | 32.6 | 55.22 | |
| GRPOBackbone=Qwen3-4B2025.12 | 32.19 | 55.08 | |
| RLOOBackbone=Qwen3-4B2025.12 | 30.83 | 54 | |
| VanillaBackbone=Qwen3-8B2025.12 | 26.98 | 48.49 | |
| VanillaBackbone=Qwen3-4B2025.12 | 23.2 | 47.44 | |
| Reinforce++Backbone=Qwen3-4B2025.12 | 17.4 | 45.03 | |
| Reinforce++Backbone=Llama-OctoThinker-8B-Base2025.12 | 7.72 | 26.43 | |
| BuPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 4.69 | 27.79 | |
| RLOOBackbone=Llama-OctoThinker-8B-Base2025.12 | 3.54 | 22.18 | |
| GRPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 2.5 | 24.11 | |
| RLOOBackbone=Llama-OctoThinker-3B-Base2025.12 | 2.19 | 17.84 | |
| PPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 1.56 | 22.82 | |
| PPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 1.04 | 16.69 | |
| GRPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 0.63 | 18.58 | |
| BuPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 0.63 | 19.59 | |
| VanillaBackbone=Llama-OctoThinker-8B-Base2025.12 | 0.52 | 3.75 | |
| VanillaBackbone=Llama-OctoThinker-3B-Base2025.12 | 0.21 | 1.68 | |
| Reinforce++Backbone=Llama-OctoThinker-3B-Base2025.12 | 0 | 5.27 |