Mathematical Reasoning on Average (AIME24, AIME25, AMC23, MATH500, APEX)
48.9Pass@1NUDGERL
Evaluation Results
| Method | Links | |
|---|---|---|
| NUDGERLModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 48.9 | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=322026.05 | 48.7 | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=162026.05 | 47 | |
| POPEModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 46.7 | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=82026.05 | 45.4 | |
| GRPOModel=Qwen3-4B-Instruct, Rollouts (N)=642026.05 | 45.1 | |
| Base modelModel=Qwen3-4B-Instruct, Rollouts (N)=–2026.05 | 40.2 | |
| NUDGERLModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 28.5 | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=322026.05 | 28.1 | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=162026.05 | 27.9 | |
| POPEModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 27.9 | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=82026.05 | 26.8 | |
| Base modelModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=–2026.05 | 22.5 | |
| GRPOModel=Olmo3-7B-Instruct-SFT, Rollouts (N)=642026.05 | 16 |