Mathematical Reasoning on MATH500 (val)
87.64AccuracyRLOO
Evaluation Results
| Method | Links | |
|---|---|---|
| RLOOAlgorithm=RLOO, Base Model=Qwen3-8B-Base, Sampled generations=16, RL Setup=Zero RL2026.04 | 87.64 | |
| VC-PPOAlgorithm=VC-PPO, Base Model=Qwen3-8B-Base, Sampled generations=16, RL Setup=Zero RL2026.04 | 87.54 | |
| GenACAlgorithm=GenAC, Base Model=Qwen3-8B-Base, Sampled generations=16, RL Setup=Zero RL2026.04 | 87.48 | |
| GRPOAlgorithm=GRPO, Base Model=Qwen3-8B-Base, Sampled generations=16, RL Setup=Zero RL2026.04 | 87.2 | |
| SyncLag (k)=0, Steps=400, GPU hours=134.42026.02 | 72 | |
| VCPOLag (k)=10, Steps=400, GPU hours=92.82026.02 | 71.6 | |
| OBLR-POModel=Qwen3-8B-Base2025.11 | 70.4 | |
| ReMaxModel=Qwen3-8B-Base2025.11 | 69.8 | |
| GRPOModel=Qwen3-8B-Base2025.11 | 69.6 | |
| RLOOModel=Qwen3-8B-Base2025.11 | 68.8 | |
| RLOOModel=Qwen3-4B-Base2025.11 | 67.8 | |
| OBLR-POModel=Qwen3-4B-Base2025.11 | 67.8 | |
| GRPOModel=Qwen3-4B-Base2025.11 | 67.6 | |
| ReMaxModel=Qwen3-4B-Base2025.11 | 65.6 | |
| PPOModel=Qwen3-8B-Base2025.11 | 64.6 | |
| PPOModel=Qwen3-4B-Base2025.11 | 59.2 | |
| Base2026.02 | 40.2 |