Mathematical Reasoning on AMC (Pass@N and Token Usage)
56.8Mean@16OptPO-GRPO
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| OptPO-GRPOBackbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 56.8 | 77.1 | 57.8 | — | 49.06 | |
| OptPO+Rei++Backbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 56 | 79.5 | 57.8 | — | 43.5 | |
| TTRL-PPOBackbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=PPO2025.12 | 54.3 | 79.5 | 55.4 | — | — | |
| OptPO-PPOBackbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=PPO2025.12 | 53.2 | 79.5 | 55.4 | — | 41.66 | |
| TTRL-Rei++Backbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 52.9 | 80.7 | 55.4 | — | — | |
| TTRL-GRPOBackbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 52.1 | 78.3 | 54.2 | — | — | |
| OptPO-GRPOBackbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 48.5 | 81.9 | 54.2 | — | 26.93 | |
| TTRL-PPOBackbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=PPO2025.12 | 48.1 | 83.1 | 55.4 | — | — | |
| OptPO-PPOBackbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=PPO2025.12 | 48.1 | 83.1 | 53 | — | 26.8 | |
| TTRL-GRPOBackbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 47.7 | 83.1 | 54.2 | — | — | |
| TTRL-Rei++Backbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 39.3 | 81.9 | 49.4 | — | — | |
| OptPO-Rei++Backbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 38.3 | 81.9 | 44.6 | — | 27.42 | |
| OptPO-PPOBackbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=PPO2025.12 | 17.1 | 54.2 | 22.9 | — | 41.76 | |
| OptPO-GRPOBackbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 15.2 | 55.4 | 16.9 | — | 78.36 | |
| TTRL-GRPOBackbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 14.9 | 51.8 | 18.1 | — | — | |
| TTRL-PPOBackbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=PPO2025.12 | 13.9 | 55.4 | 19.3 | — | — | |
| TTRL-Rei++Backbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 11.4 | 54.2 | 16.9 | — | — | |
| OptPO-Rei++Backbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 11 | 51.8 | 13.3 | — | 18.96 |