Automated Theorem Proving on MiniF2F (test)
99.6Success RateSeed-Prover
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Seed-ProverMethod=Whole-proof, Model Size=unknown, Sample Budget=unknown2025.06 | 99.6 | — | — | — | |
| Delta-Prover (w/ Gemini 2.5 Pro)Method=Agent, Model Size=unknown, Sample Budget=163842025.06 | 95.9 | — | — | — | |
| Goedel-Prover-V2Method=Whole-proof, Model Size=32B, Sample Budget=81922025.06 | 92.2 | — | — | — | |
| Goedel-Prover-V2Method=Whole-proof, Model Size=32B, Sample Budget=10242025.06 | 91.8 | — | — | — | |
| DeepSeek-Prover-V2Method=Whole-proof, Model Size=671B, Sample Budget=81922025.06 | 88.9 | — | — | — | |
| Goedel-Prover-V2Method=Whole-proof, Model Size=32B, Sample Budget=322025.06 | 88.1 | — | — | — | |
| Prover Agent w/ Ensemble of Goedel-Prover-V2 and DeepSeek-Prover-V2Method=Agent (Final proof synthesis w/ lemma), Model Size=8B, Sample Budget=2602025.06 | 88.1 | — | — | — | |
| Prover Agent w/ Ensemble of Goedel-Prover-V2 and DeepSeek-Prover-V2Method=Agent (Direct proving w/ iterative refinement), Model Size=8B, Sample Budget=1002025.06 | 86.9 | — | — | — | |
| DeepSeek-Prover-V2Method=Whole-proof, Model Size=671B, Sample Budget=10242025.06 | 86.6 | — | — | — | |
| Prover Agent w/ Goedel-Prover-V2Method=Agent (Final proof synthesis w/ lemma), Model Size=8B, Sample Budget=2602025.06 | 86.5 | — | — | — | |
| Goedel-Prover-V2 (Small)Method=Whole-proof, Model Size=8B, Sample Budget=5122025.06 | 85.7 | — | — | — | |
| Prover Agent w/ Goedel-Prover-V2Method=Agent (Direct proving w/ iterative refinement), Model Size=8B, Sample Budget=1002025.06 | 85.7 | — | — | — | |
| Prover Agent w/ Ensemble of Goedel-Prover-V2 and DeepSeek-Prover-V2Method=Agent (Direct proving w/o iterative refinement), Model Size=8B, Sample Budget=502025.06 | 85.7 | — | — | — | |
| Goedel-Prover-V2 (Small)Method=Whole-proof, Model Size=8B, Sample Budget=2562025.06 | 85.2 | — | — | — | |
| Prover Agent w/ Goedel-Prover-V2Method=Agent (Direct proving w/o iterative refinement), Model Size=8B, Sample Budget=502025.06 | 84.4 | — | — | — | |
| Goedel-Prover-V2 (Small)Method=Whole-proof, Model Size=8B, Sample Budget=642025.06 | 83.3 | — | — | — | |
| Prover Agent w/ DeepSeek-Prover-V2Method=Agent (Final proof synthesis w/ lemma), Model Size=8B, Sample Budget=2602025.06 | 82.8 | — | — | — | |
| DeepSeek-Prover-V2 (Small)Method=Whole-proof, Model Size=7B, Sample Budget=81922025.06 | 82 | — | — | — | |
| Prover Agent w/ DeepSeek-Prover-V2Method=Agent (Direct proving w/ iterative refinement), Model Size=8B, Sample Budget=1002025.06 | 82 | — | — | — | |
| DSP+ (w/ DeepSeek-R1, DeepSeek-V3, and BFS-Prover)Method=Informal + Tree search, Model Size=671B, Sample Budget=10242025.06 | 80.7 | — | — | — | |
| Kimina-Prover-PreviewMethod=Whole-proof, Model Size=72B, Sample Budget=81922025.06 | 80.7 | — | — | — | |
| DeepSeek-Prover-V2 (Small)Method=Whole-proof, Model Size=7B, Sample Budget=10242025.06 | 79.9 | — | — | — | |
| Prover Agent w/ DeepSeek-Prover-V2Method=Agent (Direct proving w/o iterative refinement), Model Size=8B, Sample Budget=502025.06 | 79.9 | — | — | — | |
| DSP+ (w/ QwQ, DeepSeek-V3, and BFS-Prover)Method=Informal + Tree search, Model Size=671B, Sample Budget=10242025.06 | 79.5 | — | — | — | |
| Leanabell-Prover-V2-DSMethod=Whole-proof, Model Size=7B, Sample Budget=1282025.06 | 78.2 | — | — | — | |
| Kimina-Prover-PreviewMethod=Whole-proof, Model Size=72B, Sample Budget=10242025.06 | 77.9 | — | — | — | |
| Leanabell-Prover-V2-DSMethod=Whole-proof, Model Size=7B, Sample Budget=322025.06 | 76.6 | — | — | — | |
| DeepSeek-Prover-V2 (Small)Method=Whole-proof, Model Size=7B, Sample Budget=322025.06 | 75.6 | — | — | — | |
| DSP+ (w/ QwQ, DeepSeek-V3, and BFS-Prover)Method=Informal + Tree search, Model Size=671B, Sample Budget=1282025.06 | 74.2 | — | — | — | |
| BFS-ProverCritic=No, Search=BFS, Tactic Budget=accumulative2025.02 | 72.95 | — | — | — | |
| BFS-ProverMethod=Tree search, Model Size=7B, Sample Budget=2048 × 2 × 6002025.06 | 70.8 | — | — | — | |
| Kimina-Prover-Preview-DistillMethod=Whole-proof, Model Size=7B, Sample Budget=10242025.06 | 70.8 | — | — | — | |
| DeepSeek-Prover-V2Size=7B, Training / Inference Setting=AR SFT + RLVR + Long CoT2026.06 | 70.49 | — | — | — | |
| Leanabell-Prover-V2-KMMethod=Whole-proof, Model Size=7B, Sample Budget=1282025.06 | 70.4 | — | — | — | |
| HunyuanProverCritic=Yes, Search=BFS, Tactic Budget=600 × 8 × 4002025.02 | 68.4 | — | — | — | |
| HunyuanProver v16 + BFS + DCMethod=Tree search, Model Size=7B, Sample Budget=600 × 8 × 4002025.06 | 68.4 | — | — | — | |
| Leanabell-Prover-V2-KMMethod=Whole-proof, Model Size=7B, Sample Budget=322025.06 | 68.4 | — | — | — | |
| STPMethod=Whole-proof, Model Size=7B, Sample Budget=256002025.06 | 67.6 | — | — | — | |
| ProofAug + Mix. + Dataset CurationModel=deepseek-math-7b-base, Sample Budget=2100†, Proof Assistant=Isabelle, Evaluation Strategy=mixed strategy, Additional Config=Dataset Curation2025.01 | 66 | — | — | — | |
| InternLM2.5-StepProverCritic=Yes, Search=BFS, Tactic Budget=256 × 32 × 6002025.02 | 65.9 | — | — | — | |
| BFS + CGModel=InternLM2.5-StepProver (7B), Sample Budget=256 x 32 x 600, Proof Assistant=Lean2025.01 | 65.9 | — | — | — | |
| InternLM2.5-StepProver-BF + CGMethod=Tree search, Model Size=7B, Sample Budget=256 × 32 × 6002025.06 | 65.9 | — | — | — | |
| Goedel-Prover-SFTMethod=Whole-proof, Model Size=7B, Sample Budget=256002025.06 | 64.7 | — | — | — | |
| Prover Agent w/ Goedel-Prover-V2Method=Agent (Direct proving w/o iterative refinement), Model Size=8B, Sample Budget=12025.06 | 64.3 | — | — | — | |
| Prover Agent w/ Ensemble of Goedel-Prover-V2 and DeepSeek-Prover-V2Method=Agent (Direct proving w/o iterative refinement), Model Size=8B, Sample Budget=12025.06 | 64.3 | — | — | — | |
| DeepSeek-Prover-V1.5Critic=No, Search=MCTS, Tactic Budget=32 × 16 × 4002025.02 | 63.5 | — | — | — | |
| RMaxTSModel=DeepSeek-Prover-V1.5-RL (7B), Sample Budget=32 x 6400†, Proof Assistant=Lean, Evaluation Strategy=mixed strategy2025.01 | 63.5 | — | — | — | |
| DeepSeek-Prover-V1.5-RL + RMaxTSMethod=Tree search, Model Size=7B, Sample Budget=32 × 16 × 4002025.06 | 63.5 | — | — | — | |
| Kimina-Prover-Preview-DistillMethod=Whole-proof, Model Size=7B, Sample Budget=322025.06 | 63.1 | — | — | — | |
| ProofAug + Mixed StrategyModel=deepseek-math-7b-base, Sample Budget=1400†, Proof Assistant=Isabelle, Evaluation Strategy=mixed strategy2025.01 | 61.9 | — | — | — | |
| DeepSeek-Prover-V2Method=Whole-proof, Model Size=671B, Sample Budget=12025.06 | 61.9 | — | — | — | |
| Prover Agent w/ DeepSeek-Prover-V2Method=Agent (Direct proving w/o iterative refinement), Model Size=8B, Sample Budget=12025.06 | 61.5 | — | — | — | |
| Leanabell-Prover-GD-RLMethod=Whole-proof, Model Size=7B, Sample Budget=1282025.06 | 61.1 | — | — | — | |
| Goedel-Prover-V2 (Small)Method=Whole-proof, Model Size=8B, Sample Budget=12025.06 | 60.8 | — | — | — | |
| DeepSeek-Prover-V2 (Small)Method=Whole-proof, Model Size=7B, Sample Budget=12025.06 | 58.6 | — | — | — | |
| Kimina-Prover-PreviewSize=1.5B, Training / Inference Setting=AR SFT + RLVR + Long CoT2026.06 | 56.2 | — | — | — | |
| Subgoal-XLModel=Fine-tuned Llama-8B, Sample Budget=16384†, Proof Assistant=Isabelle, Evaluation Strategy=mixed strategy2025.01 | 56.1 | — | — | — | |
| ProofAug (0-shot) + ERPModel=deepseek-math-7b-base, Sample Budget=500, Proof Assistant=Isabelle, Prompting Strategy=0-shot, Module=ERP2025.01 | 56.1 | — | — | — | |
| ProofAug (0-shot)Model=deepseek-math-7b-base, Sample Budget=500, Proof Assistant=Isabelle, Prompting Strategy=0-shot2025.01 | 54.5 | — | — | — | |
| Kimina-Prover-PreviewMethod=Whole-proof, Model Size=72B, Sample Budget=12025.06 | 52.9 | — | — | — | |
| ProofAugModel=deepseek-math-7b-base, Sample Budget=100, Proof Assistant=Isabelle2025.01 | 52.5 | — | — | — | |
| DSP+ (w/ QwQ, DeepSeek-V3, and BFS-Prover)Method=Informal + Tree search, Model Size=671B, Sample Budget=12025.06 | 52.5 | — | — | — | |
| Kimina-Prover-Preview-DistillMethod=Whole-proof, Model Size=7B, Sample Budget=12025.06 | 52.5 | — | — | — | |
| LyraModel=GPT-4, Sample Budget=200, Proof Assistant=Isabelle2025.01 | 51.2 | — | — | — | |
| LEGO-ProverModel=mixed GPTs, Sample Budget=100, Proof Assistant=Isabelle2025.01 | 50 | — | — | — | |
| DeepSeek-Prover-V1.5Size=7B, Training / Inference Setting=AR SFT + RL2026.06 | 50 | — | — | — | |
| Diffusion-ProofSize=7B, Training / Inference Setting=dLLM SFT2026.06 | 50 | — | — | — | |
| DSP baselineModel=deepseek-math-7b-base, Sample Budget=100, Proof Assistant=Isabelle2025.01 | 49.2 | — | — | — | |
| LyraModel=GPT-4, Sample Budget=100, Proof Assistant=Isabelle2025.01 | 47.1 | — | — | — | |
| Lean-STaRSize=7B, Training / Inference Setting=AR SFT / expert iteration2026.06 | 46.3 | — | — | — | |
| DeepSeek-Prover-V1Size=7B, Training / Inference Setting=AR SFT2026.06 | 46.3 | — | — | — | |
| ProofAugModel=deepseek-math-7b-base, Sample Budget=10, Proof Assistant=Isabelle2025.01 | 44.7 | — | — | — | |
| POETRYModel=Fine-tuned ProofGPT (1.3B), Sample Budget=1 x 32 x 128, Proof Assistant=Isabelle2025.01 | 42.2 | — | — | — | |
| HTPSModel=Evariste (600M), Sample Budget=64 x 5000, Proof Assistant=Lean2025.01 | 41 | — | — | — | |
| DSP baselineModel=deepseek-math-7b-base, Sample Budget=10, Proof Assistant=Isabelle2025.01 | 40.6 | — | — | — | |
| DSPModel=CodeX, Sample Budget=100, Proof Assistant=Isabelle2025.01 | 39.3 | — | — | — | |
| Subgoal-XLModel=Fine-tuned Llama-8B, Sample Budget=64, Proof Assistant=Isabelle2025.01 | 39.3 | — | — | — | |
| ProofAugModel=deepseek-math-7b-base, Sample Budget=1, Proof Assistant=Isabelle2025.01 | 36.5 | — | — | — | |
| TheoremLlamaSize=8B, Training / Inference Setting=AR SFT2026.06 | 35.7 | — | — | — | |
| Thorsampling temperature=1.22022.05 | 29.9 | — | — | — | |
| Expert iteration2022.05 | 29.6 | — | — | — | |
| DSP baselineModel=deepseek-math-7b-base, Sample Budget=1, Proof Assistant=Isabelle2025.01 | 28.7 | — | — | — | |
| Hierarchical AttentionK=64, Search strategy=single-pass sampling (SPS)2025.04 | 27.87 | 1.85 | 0.93 | 23.21 | |
| Language model ∪ Sledgehammer2022.05 | 27.5 | — | — | — | |
| Hierarchical AttentionK=32, Search strategy=single-pass sampling (SPS)2025.04 | 26.64 | 1.78 | 0.97 | 15.38 | |
| Hierarchical AttentionK=16, Search strategy=single-pass sampling (SPS)2025.04 | 26.23 | 1.92 | 1.04 | 26 | |
| Hierarchical AttentionK=8, Search strategy=single-pass sampling (SPS)2025.04 | 25 | 1.86 | 0.95 | 51.16 | |
| PACT2022.05 | 24.6 | — | — | — | |
| Language model2022.05 | 24.2 | — | — | — | |
| Hierarchical AttentionK=4, Search strategy=single-pass sampling (SPS)2025.04 | 24.18 | — | — | — | |
| BaselineK=32, Search strategy=single-pass sampling (SPS)2025.04 | 23.36 | 1.83 | — | — | |
| BaselineK=64, Search strategy=single-pass sampling (SPS)2025.04 | 23.36 | 2 | — | — | |
| Hierarchical AttentionK=2, Search strategy=single-pass sampling (SPS)2025.04 | 20.9 | — | — | — | |
| BaselineK=16, Search strategy=single-pass sampling (SPS)2025.04 | 20.49 | 1.85 | — | — | |
| BaselineK=8, Search strategy=single-pass sampling (SPS)2025.04 | 19.63 | 1.95 | — | — | |
| Hierarchical AttentionK=1, Search strategy=single-pass sampling (SPS)2025.04 | 18.44 | — | — | — | |
| BaselineK=4, Search strategy=single-pass sampling (SPS)2025.04 | 16.8 | — | — | — | |
| BaselineK=2, Search strategy=single-pass sampling (SPS)2025.04 | 12.3 | — | — | — | |
| Sledgehammer2022.05 | 10.4 | — | — | — | |
| BaselineK=1, Search strategy=single-pass sampling (SPS)2025.04 | 9.84 | — | — | — |