Mathematical Reasoning on MATH500 (Avg@16 and Average)
88.05Avg@16 ScoreGRPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPOBackbone=Qwen3-8B2025.12 | 88.05 | 64.23 | |
| BuPOBackbone=Qwen3-8B2025.12 | 87.76 | 66.36 | |
| RLOOBackbone=Qwen3-8B2025.12 | 87.32 | 63.36 | |
| PPOBackbone=Qwen3-8B2025.12 | 86.2 | 58.41 | |
| Reinforce++Backbone=Qwen3-8B2025.12 | 86.05 | 60.41 | |
| BuPOBackbone=Qwen3-4B2025.12 | 84.9 | 58.51 | |
| PPOBackbone=Qwen3-4B2025.12 | 83.64 | 55.22 | |
| RLOOBackbone=Qwen3-4B2025.12 | 82.73 | 54 | |
| GRPOBackbone=Qwen3-4B2025.12 | 82.41 | 55.08 | |
| Reinforce++Backbone=Qwen3-4B2025.12 | 80.63 | 45.03 | |
| VanillaBackbone=Qwen3-8B2025.12 | 80.46 | 48.49 | |
| VanillaBackbone=Qwen3-4B2025.12 | 80.29 | 47.44 | |
| BuPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 62.05 | 27.79 | |
| Reinforce++Backbone=Llama-OctoThinker-8B-Base2025.12 | 59.55 | 26.43 | |
| PPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 56.97 | 22.82 | |
| GRPOBackbone=Llama-OctoThinker-8B-Base2025.12 | 56.89 | 24.11 | |
| RLOOBackbone=Llama-OctoThinker-8B-Base2025.12 | 55.97 | 22.18 | |
| BuPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 49.79 | 19.59 | |
| GRPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 46.07 | 18.58 | |
| PPOBackbone=Llama-OctoThinker-3B-Base2025.12 | 43.23 | 16.69 | |
| RLOOBackbone=Llama-OctoThinker-3B-Base2025.12 | 41.93 | 17.84 | |
| Reinforce++Backbone=Llama-OctoThinker-3B-Base2025.12 | 11.59 | 5.27 | |
| VanillaBackbone=Llama-OctoThinker-8B-Base2025.12 | 9.84 | 3.75 | |
| VanillaBackbone=Llama-OctoThinker-3B-Base2025.12 | 5.26 | 1.68 |