Question Answering on GPQA (Mean@16, Pass@16, Pass@1, Toks.)
37.9Mean@16TTRL-PPO
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TTRL-PPOBackbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=PPO2025.12 | 37.9 | 83.8 | 40.1 | — | — | |
| OptPO-PPOBackbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=PPO2025.12 | 36.8 | 88.8 | 39.1 | — | 39.83 | |
| OptPO-GRPOBackbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 35.7 | 84.8 | 38.6 | — | 40.31 | |
| TTRL-GRPOBackbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 35.1 | 85.3 | 39.1 | — | — | |
| TTRL-Rei++Backbone=Qwen2.5-7B, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 33 | 85.3 | 32.5 | — | — | |
| OptPO+Rei++Backbone=Qwen2.5-7B, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 33 | 84.3 | 34 | — | 45.47 | |
| TTRL-Rei++Backbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 27.1 | 85.8 | 30.5 | — | — | |
| OptPO-GRPOBackbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 26.8 | 87.8 | 34.5 | — | 43.2 | |
| OptPO-PPOBackbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=PPO2025.12 | 26.7 | 81.7 | 31 | — | 46.19 | |
| OptPO-Rei++Backbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 26.7 | 85.8 | 28.9 | — | 41.96 | |
| OptPO-PPOBackbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=PPO2025.12 | 26.6 | 89.8 | 32.5 | — | 45.06 | |
| TTRL-GRPOBackbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 26.5 | 87.8 | 30.5 | — | — | |
| OptPO-Rei++Backbone=Qwen2.5-Math-1.5B, Methodology=OptPO, RL Algorithm=Rei++2025.12 | 26.5 | 90.9 | 28.9 | — | 43.52 | |
| TTRL-GRPOBackbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=GRPO2025.12 | 26.4 | 82.2 | 28.4 | — | — | |
| TTRL-Rei++Backbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=Rei++2025.12 | 26.3 | 93.9 | 32 | — | — | |
| TTRL-PPOBackbone=Qwen2.5-Math-1.5B, Methodology=TTRL, RL Algorithm=PPO2025.12 | 25.9 | 87.8 | 29.9 | — | — | |
| TTRL-PPOBackbone=Llama-3.2-1B-Instruct, Methodology=TTRL, RL Algorithm=PPO2025.12 | 25.7 | 83.2 | 28.9 | — | — | |
| OptPO-GRPOBackbone=Llama-3.2-1B-Instruct, Methodology=OptPO, RL Algorithm=GRPO2025.12 | 25.2 | 81.2 | 26.9 | — | 44.74 |