Mathematical Reasoning on AIME 2024 (Pass@16, Mean@16, Token Metrics)
24.4Mean Score @16TTRL-PPO
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TTRL-PPOBackbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=PPO2025.12 | 24.4 | 46.7 | 23.3 | — | — | |
| OptPO-PPOBackbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=PPO2025.12 | 24 | 43.3 | 26.7 | — | 42.85 | |
| TTRL-GRPOBackbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 23.5 | 40 | 23.3 | — | — | |
| OptPO-GRPOBackbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 23.5 | 43.3 | 23.3 | — | 38.56 | |
| TTRL-Rei++Backbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 20.8 | 40 | 23.3 | — | — | |
| OptPO+Rei++Backbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 20.6 | 40 | 23.3 | — | 41.45 | |
| TTRL-PPOBackbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=PPO2025.12 | 18.5 | 46.7 | 23.3 | — | — | |
| OptPO-PPOBackbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=PPO2025.12 | 16 | 46.7 | 23.3 | — | 15.2 | |
| TTRL-GRPOBackbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 15.6 | 43.3 | 20 | — | — | |
| OptPO-GRPOBackbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 14.6 | 46.7 | 20 | — | 17.25 | |
| TTRL-Rei++Backbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 11.9 | 43.3 | 20 | — | — | |
| OptPO-Rei++Backbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 10.4 | 46.7 | 20 | — | 14.35 | |
| TTRL-PPOBackbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=PPO2025.12 | 3.8 | 30 | 6.7 | — | — | |
| OptPO-PPOBackbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=PPO2025.12 | 3.8 | 30 | 6.7 | — | 13.7 | |
| TTRL-GRPOBackbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 3.5 | 23.3 | 6.7 | — | — | |
| OptPO-GRPOBackbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 1.7 | 20 | 6.7 | — | 47.62 | |
| OptPO-Rei++Backbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 1.7 | 23.3 | 6.7 | — | 30.02 | |
| TTRL-Rei++Backbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 1.5 | 16.7 | 6.7 | — | — |