Theorem Proving on miniF2F Lean (val)
60.2Cumulative Pass RateDeepSeekMath-Base
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeekMath-BaseModel size=7B, Generation times=cumulative2024.05 | 60.2 | — | |
| EvaristeOnline training statements=miniF2F-valid, Train time (A100 days)=13602022.05 | 58.6 | — | |
| Curriculum LearningModel size=837M, Generation times=cumulative2024.05 | 58.6 | — | |
| Curriculum LearningModel size=837M, Generation times=64 x 8 x 5122024.05 | 47.3 | — | |
| Curriculum LearningModel size=837M, Generation times=8 x 8 x 5122024.05 | 41.2 | — | |
| Curriculum LearningModel size=837M, Generation times=1 x 8 x 5122024.05 | 33.6 | — | |
| Proof Artifact Co-TrainingModel size=837M, Generation times=8 x 8 x 5122024.05 | 29.3 | — | |
| GPT-4-turbo 0409Generation times=642024.05 | 25.4 | — | |
| DeepSeekMath-BaseModel size=7B, Generation times=1282024.05 | 25.4 | — | |
| Proof Artifact Co-TrainingModel size=837M, Generation times=1 x 8 x 5122024.05 | 23.9 | — | |
| Evariste-1dOnline training statements=miniF2F-curriculum, Train time (A100 days)=2302022.05 | — | 46.7 | |
| Evariste-7dOnline training statements=miniF2F-curriculum, Train time (A100 days)=16202022.05 | — | 47.5 | |
| GPT-fTrain time (A100 days)=20002022.05 | — | 47.3 | |
| SupervisedTrain time (A100 days)=502022.05 | — | 38.5 |