Formal Theorem Proving on ProofNet
24.26AccuracyGOEDEL VALUE
Evaluation Results
| Method | Links | |
|---|---|---|
| GOEDEL VALUESampling Budget=64, Search Strategy=Value-guided tree search2026.03 | 24.26 | |
| GOEDEL RANDOMSampling Budget=64, Search Strategy=Random search2026.03 | 23.72 | |
| GOEDEL DIRECTSampling Budget=64, Supervision Strategy=direct synthesis2026.03 | 19.14 | |
| KIMINA RANDOMSampling Budget=64, Search Strategy=Random search2026.03 | 15.63 | |
| GOEDEL-V2Sampling Budget=642026.03 | 15.63 | |
| KIMINA VALUESampling Budget=64, Search Strategy=Value-guided tree search2026.03 | 15.36 | |
| KIMINA DIRECTSampling Budget=64, Supervision Strategy=direct synthesis2026.03 | 14.82 | |
| KIMINASampling Budget=642026.03 | 14.56 | |
| KG-ProverLLM Model=o1-mini, Max Attempts=32025.02 | 6.99 | |
| KG-ProverLLM Model=GPT 4o, Max Attempts=32025.02 | 6.45 | |
| RAGLLM Model=o1-mini, Max Attempts=32025.02 | 5.91 | |
| RAGLLM Model=GPT 4o, Max Attempts=32025.02 | 5.38 | |
| KG-ProverLLM Model=Deepseek R1, Max Attempts=32025.02 | 5.38 | |
| KG-ProverLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 4.84 | |
| KG-ProverLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 4.3 | |
| KG-ProverLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 4.3 | |
| BaseLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 3.76 | |
| BaseLLM Model=o1-mini, Max Attempts=32025.02 | 3.76 | |
| RAGLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 3.76 | |
| RAGLLM Model=Deepseek R1, Max Attempts=32025.02 | 3.76 | |
| RAGLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 3.76 | |
| RAGLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 3.76 | |
| BaseLLM Model=GPT 4o, Max Attempts=32025.02 | 3.23 | |
| BaseLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 2.69 | |
| BaseLLM Model=Deepseek R1, Max Attempts=32025.02 | 2.69 | |
| BaseLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 2.15 |