Mathematical Reasoning on Mathematical Reasoning Aggregate
75.4Average ScoreOLMo-32B-Instruct (Teacher)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| OLMo-32B-Instruct (Teacher)Model Category=Off-the-shelf2026.06 | 75.4 | — | — | |
| SEAD (full)Model Category=Proposed (SEAD components)2026.06 | 64 | — | — | |
| Token Zones + KL AnnealingModel Category=Proposed (SEAD components)2026.06 | 63.7 | — | — | |
| Token Zones onlyModel Category=Proposed (SEAD components)2026.06 | 59.5 | — | — | |
| KL Annealing onlyModel Category=Proposed (SEAD components)2026.06 | 59.4 | — | — | |
| OPD (RKL k=1)Model Category=Baseline2026.06 | 59.2 | — | — | |
| RLADModel=Qwen2.5-1.5B-DS, Context=8K2026.02 | 59.1 | — | — | |
| OLMo-7B-Instruct (Student)Model Category=Off-the-shelf2026.06 | 58.2 | — | — | |
| OPSDModel Category=Baseline2026.06 | 58.2 | — | — | |
| GRPOModel Category=Baseline2026.06 | 58 | — | — | |
| Qwen2.5-1.5B-DSTraining=KDRL, Context=8K2026.02 | 57.9 | — | — | |
| RLADModel=Qwen3-1.7B, Context=8K2026.02 | 57.8 | — | — | |
| Qwen2.5-1.5B-DSTraining=GRPO, Context=8K2026.02 | 56.5 | — | — | |
| Qwen3-1.7BTraining=KDRL, Context=8K2026.02 | 55.1 | — | — | |
| Qwen3-1.7BTraining=GRPO, Context=8K2026.02 | 51.2 | — | — | |
| GRPO + CBS*Backbone=Qwen3-8B-Base2026.02 | 46.63 | — | 5.2 | |
| DAPO + CBS*Backbone=Qwen3-8B-Base2026.02 | 45.63 | — | -1.7 | |
| Qwen2.5-1.5B-DSTraining=-, Context=8K2026.02 | 44.7 | — | — | |
| GSPO + CBSBackbone=Qwen3-4B-Base2026.02 | 43.92 | — | -0.4 | |
| GRPO-SGBase Model=Qwen2.5-7B, Training Recipe=DSR (DeepScaleR)2025.10 | 43.04 | — | 6 | |
| GRPO + CBSBackbone=Qwen3-4B-Base2026.02 | 42.41 | — | 2.9 | |
| GRPO-SGBase Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 42.32 | — | 2.7 | |
| MathForgeBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 42.17 | 4.56 | — | |
| Qwen3-1.7BTraining=-, Context=8K2026.02 | 42.1 | — | — | |
| ARBase Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 41.72 | — | 1.2 | |
| 80/20Base Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 41.57 | — | 0.9 | |
| LoptiBase Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 41.26 | — | 0.1 | |
| GRPOBase Model=Qwen2.5-7B, Training Recipe=ORZ (Open Reasoner-Zero)2025.10 | 41.22 | — | — | |
| MQRBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 41.04 | 3.43 | — | |
| LoptiBase Model=Qwen2.5-7B, Training Recipe=DSR (DeepScaleR)2025.10 | 40.8 | — | 0.5 | |
| GRPOBase Model=Qwen2.5-7B, Training Recipe=DSR (DeepScaleR)2025.10 | 40.6 | — | — | |
| ARBase Model=Qwen2.5-7B, Training Recipe=DSR (DeepScaleR)2025.10 | 40.38 | — | 0.5 | |
| 80/20Base Model=Qwen2.5-7B, Training Recipe=DSR (DeepScaleR)2025.10 | 39.89 | — | 1.7 | |
| DGPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 39.79 | 2.18 | — | |
| GRPO-ADBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 38.26 | 0.65 | — | |
| DAPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 37.94 | 0.33 | — | |
| GPGBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 37.93 | 0.32 | — | |
| GSPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 37.71 | 0.1 | — | |
| GRPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 37.61 | — | — | |
| Dr.GRPOBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 37.4 | -0.21 | — | |
| GSPO + CBS*Backbone=Qwen3-1.7B-Base2026.02 | 32.41 | — | 5 | |
| GRPO + CBS*Backbone=Qwen3-1.7B-Base2026.02 | 31.08 | — | 7.5 | |
| GRPO + CBSBackbone=Qwen3-1.7B-Base2026.02 | 30.17 | — | 4.4 | |
| DAPO + CBSBackbone=Qwen3-1.7B-Base2026.02 | 30 | — | 15.7 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B2025.10 | 27.66 | — | — | |
| Base ModelBackbone=Qwen2.5-Math-7B, Training Dataset=MATH, Evaluation Protocol=Zero-shot2026.01 | 22.04 | — | — |