Machine Learning Engineering on MLE-bench (held-out task instances)
58.6Accuracy (%)Full ExIt
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Full ExItBackbone=DeepSeek-R1-Distill-Qwen-7B, Improvement steps (K)=16, Number of training runs=32025.09 | 58.6 | 8.4 | |
| Diverge (ExIt ablation)Backbone=DeepSeek-R1-Distill-Qwen-7B, Improvement steps (K)=16, Number of training runs=32025.09 | 57.3 | 10.1 | |
| GRPO + curriculumBackbone=DeepSeek-R1-Distill-Qwen-7B, Improvement steps (K)=16, Number of training runs=32025.09 | 53 | 11.9 | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-7B, Improvement steps (K)=16, Number of training runs=32025.09 | 48 | 9.1 | |
| Improve (ExIt ablation)Backbone=DeepSeek-R1-Distill-Qwen-7B, Improvement steps (K)=16, Number of training runs=32025.09 | 47.8 | 9.4 | |
| Base modelBackbone=DeepSeek-R1-Distill-Qwen-7B, Improvement steps (K)=16, Number of training runs=32025.09 | 4.2 | 2.4 |