Mathematical Reasoning on APEX Shortlist (pass@1)
5.8pass@1GRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=322026.05 | 5.8 | |
| NUDGERLModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 5.3 | |
| POPEModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 4.8 | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=162026.05 | 4.5 | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 4 | |
| Base modelModel=Qwen3-4B-Instruct, Rollouts (N)=–2026.05 | 3.6 | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=642026.05 | 2.7 | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=642026.05 | 2.7 | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 2.5 | |
| NUDGERLModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 2.5 | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=322026.05 | 2.4 | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=162026.05 | 2.3 | |
| POPEModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 2.3 | |
| Base modelModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=–2026.05 | 2.1 |