Theorem Proving on MiniF2F (val)
63.9Success RateInternLM2-StepProver
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| InternLM2-StepProverLanguage=Lean-based, Base=InternLM2-Math-plus, Synthetic Size=1.155B Tokens, Search Budget=64 x 3200, Tree search strategy=true2024.08 | 63.9 | — | — | — | |
| SubgoalXLBase model=Llama-3-8B, Use human-written informal proofs=true2024.08 | 61.9 | — | — | — | |
| SubgoalXLLanguage=Isabelle-based, Base=Llama-3, Synthetic Size=38k Theorems, Search Budget=16384, Tree search strategy=true2024.08 | 61.9 | — | — | — | |
| HTPSLanguage=Lean-based, Tree search strategy=true2024.08 | 58.6 | — | — | — | |
| LEGO-ProverBase model=GPT-3.5-Turbo, Use human-written informal proofs=true2024.08 | 55.3 | — | — | — | |
| LyraBase model=GPT-4, Use human-written informal proofs=true2024.08 | 55.3 | — | — | — | |
| LEGO-ProverLanguage=Isabelle-based, Base=GPT-3.5-Turbo, Search Budget=100, Tree search strategy=true2024.08 | 55.3 | — | — | — | |
| LyraLanguage=Isabelle-based, Base=GPT-4, Search Budget=200, Human-written informal proofs integration=true2024.08 | 55.3 | — | — | — | |
| Subgoal-ProverBase model=GPT-3.5-Turbo, Use human-written informal proofs=false2024.08 | 48 | — | — | — | |
| Subgoal-ProverLanguage=Isabelle-based, Base=GPT-3.5-Turbo, Search Budget=1002024.08 | 48 | — | — | — | |
| DSPBase model=Codex, Use human-written informal proofs=true2024.08 | 42.6 | — | — | — | |
| DSPLanguage=Isabelle-based, Base=Codex, Search Budget=100, Human-written informal proofs integration=true2024.08 | 42.6 | — | — | — | |
| M2Expert iteration=22022.05 | 37.3 | — | — | — | |
| Thor + expert iterationUse human-written informal proofs=false2024.08 | 37.3 | — | — | — | |
| Thor + expert iterationLanguage=Isabelle-based, Tree search strategy=true2024.08 | 37.3 | — | — | — | |
| M1Expert iteration=12022.05 | 36.1 | — | — | — | |
| Hierarchical AttentionK=32, Search Strategy=best-first search2025.04 | 34.02 | 2 | 0.64 | 12 | |
| Hierarchical AttentionK=64, Search Strategy=best-first search2025.04 | 34.02 | 2 | 0.64 | 12.68 | |
| Expert iteration2022.05 | 33.6 | — | — | — | |
| FMSCL2022.05 | 33.6 | — | — | — | |
| Hierarchical AttentionK=16, Search Strategy=best-first search2025.04 | 33.2 | 1.89 | 0.65 | 13.89 | |
| BaselineK=32, Search Strategy=best-first search2025.04 | 31.56 | 3.11 | — | — | |
| BaselineK=64, Search Strategy=best-first search2025.04 | 31.56 | 3.11 | — | — | |
| BaselineK=16, Search Strategy=best-first search2025.04 | 31.15 | 2.89 | — | — | |
| Hierarchical AttentionK=8, Search Strategy=best-first search2025.04 | 29.51 | 2.67 | 0.94 | 9.68 | |
| Thorsampling temperature=1.22022.05 | 28.3 | — | — | — | |
| Base model (M0)2022.05 | 28.3 | — | — | — | |
| ThorUse human-written informal proofs=false2024.08 | 28.3 | — | — | — | |
| ThorLanguage=Isabelle-based, Search Budget=300, Tree search strategy=true2024.08 | 28.3 | — | — | — | |
| Language model ∪ Sledgehammer2022.05 | 27.1 | — | — | — | |
| BaselineK=8, Search Strategy=best-first search2025.04 | 27.05 | 2.83 | — | — | |
| Hierarchical AttentionK=64, Search Strategy=single-pass sampling strategy2025.04 | 26.64 | 2.38 | 1.12 | 16.33 | |
| Hierarchical AttentionK=32, Search Strategy=single-pass sampling strategy2025.04 | 25.41 | 1.92 | 1 | 27.66 | |
| Language model2022.05 | 25 | — | — | — | |
| Hierarchical AttentionK=16, Search Strategy=single-pass sampling strategy2025.04 | 24.59 | 1.89 | 0.97 | 43.18 | |
| PACT2022.05 | 23.9 | — | — | — | |
| PACT2022.05 | 23.9 | — | — | — | |
| Hierarchical AttentionK=4, Search Strategy=best-first search2025.04 | 23.77 | — | — | — | |
| BaselineK=64, Search Strategy=single-pass sampling strategy2025.04 | 21.72 | 2.12 | — | — | |
| Hierarchical AttentionK=8, Search Strategy=single-pass sampling strategy2025.04 | 21.13 | 1.73 | 0.74 | 37.5 | |
| BaselineK=4, Search Strategy=best-first search2025.04 | 20.49 | — | — | — | |
| Hierarchical AttentionK=4, Search Strategy=single-pass sampling strategy2025.04 | 20.08 | — | — | — | |
| BaselineK=32, Search Strategy=single-pass sampling strategy2025.04 | 20.08 | 1.92 | — | — | |
| BaselineK=16, Search Strategy=single-pass sampling strategy2025.04 | 18.44 | 1.95 | — | — | |
| BaselineK=8, Search Strategy=single-pass sampling strategy2025.04 | 18.03 | 2.33 | — | — | |
| Sledgehammer+heuristicUse human-written informal proofs=false2024.08 | 18 | — | — | — | |
| Sledgehammer+heuristicLanguage=Isabelle-based2024.08 | 18 | — | — | — | |
| BaselineK=4, Search Strategy=single-pass sampling strategy2025.04 | 16.8 | — | — | — | |
| Hierarchical AttentionK=2, Search Strategy=single-pass sampling strategy2025.04 | 15.16 | — | — | — | |
| BaselineK=2, Search Strategy=best-first search2025.04 | 15.16 | — | — | — | |
| Hierarchical AttentionK=1, Search Strategy=single-pass sampling strategy2025.04 | 14.75 | — | — | — | |
| Hierarchical AttentionK=2, Search Strategy=best-first search2025.04 | 14.75 | — | — | — | |
| Hierarchical AttentionK=1, Search Strategy=best-first search2025.04 | 13.52 | — | — | — | |
| BaselineK=1, Search Strategy=best-first search2025.04 | 12.7 | — | — | — | |
| BaselineK=2, Search Strategy=single-pass sampling strategy2025.04 | 12.3 | — | — | — | |
| Sledgehammer2022.05 | 9.9 | — | — | — | |
| SledgehammerUse human-written informal proofs=false2024.08 | 9.9 | — | — | — | |
| SledgehammerLanguage=Isabelle-based2024.08 | 9.9 | — | — | — | |
| BaselineK=1, Search Strategy=single-pass sampling strategy2025.04 | 9.43 | — | — | — |