Mathematical Reasoning on MATH 500 (pass@1, standard deviation)
94Pass@1 AccuracyOLMo-32B-Instruct (Teacher)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OLMo-32B-Instruct (Teacher)Model Category=Off-the-shelf, Base Model=OLMo-32B2026.06 | 94 | — | |
| SEAD (full)Model Category=Proposed (SEAD components)2026.06 | 91.2 | — | |
| Token Zones + KL AnnealingModel Category=Proposed (SEAD components)2026.06 | 91 | — | |
| KL Annealing onlyModel Category=Proposed (SEAD components)2026.06 | 87.8 | — | |
| OLMo-7B-Instruct (Student)Model Category=Off-the-shelf, Base Model=OLMo-7B2026.06 | 87.2 | — | |
| OPD (RKL k=1)Model Category=Baseline, Setting=k=12026.06 | 87 | — | |
| Token Zones onlyModel Category=Proposed (SEAD components)2026.06 | 87 | — | |
| OPSDModel Category=Baseline2026.06 | 86.4 | — | |
| GRPOModel Category=Baseline2026.06 | 84.8 | — | |
| UniRBackbone=Qwen2.5-3B, Trained Model=1.5B, Zero-shot=true, Temperature=0.62025.05 | 66.8 | — | |
| GRPO FullBackbone=Qwen2.5-3B, Trained Model=3B, Zero-shot=true, Temperature=0.62025.05 | 62.9 | — | |
| GRPO LoRABackbone=Qwen2.5-3B, Trained Model=3B, Zero-shot=true, Temperature=0.62025.05 | 60 | — | |
| Mathstral-7B-v0.1Trained Model=7B, Zero-shot=true, Temperature=0.62025.05 | 56.6 | — | |
| NuminaMath-7B-CoTTrained Model=7B, Zero-shot=true, Temperature=0.62025.05 | 55.2 | — | |
| Internlm2-math-plus-7BTrained Model=7B, Zero-shot=true, Temperature=0.62025.05 | 54.4 | — | |
| DeepSeekMath-7B-RLTrained Model=7B, Zero-shot=true, Temperature=0.62025.05 | 52.4 | — | |
| Backbone + 1.5BBackbone=Qwen2.5-3B, Zero-shot=true, Temperature=0.62025.05 | 50.6 | — | |
| UniRBackbone=Llama3.2-3B, Trained Model=1B, Zero-shot=true, Temperature=0.62025.05 | 48.8 | — | |
| Backbone onlyBackbone=Qwen2.5-3B, Zero-shot=true, Temperature=0.62025.05 | 44.2 | — | |
| GRPO FullBackbone=Llama3.2-3B, Trained Model=3B, Zero-shot=true, Temperature=0.62025.05 | 36.1 | — | |
| Backbone + 1BBackbone=Llama3.2-3B, Zero-shot=true, Temperature=0.62025.05 | 35.7 | — | |
| Backbone onlyBackbone=Llama3.2-3B, Zero-shot=true, Temperature=0.62025.05 | 33 | — | |
| GRPO LoRABackbone=Llama3.2-3B, Trained Model=3B, Zero-shot=true, Temperature=0.62025.05 | 32.1 | — |