Mathematical Reasoning on Math held-out task instances (test)
20.4AccuracyFull ExIt
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Full ExItBackbone=Llama-3.2-3B-Instruct, Improvement steps (K)=16, Number of training runs=32025.09 | 20.4 | 2 | |
| Diverge (ExIt ablation)Backbone=Llama-3.2-3B-Instruct, Improvement steps (K)=16, Number of training runs=32025.09 | 20.1 | 1.6 | |
| Improve (ExIt ablation)Backbone=Llama-3.2-3B-Instruct, Improvement steps (K)=16, Number of training runs=32025.09 | 19.6 | 1.2 | |
| GRPO + curriculumBackbone=Llama-3.2-3B-Instruct, Improvement steps (K)=16, Number of training runs=32025.09 | 18.8 | 0.9 | |
| GRPOBackbone=Llama-3.2-3B-Instruct, Improvement steps (K)=16, Number of training runs=32025.09 | 18.7 | 1.1 | |
| Base modelBackbone=Llama-3.2-3B-Instruct, Improvement steps (K)=16, Number of training runs=32025.09 | 17.4 | -0.4 |