Automated Theorem Proving on miniF2F
31.97AccuracyKG-Prover
Evaluation Results
| Method | Links | |
|---|---|---|
| KG-ProverLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 31.97 | |
| KG-ProverLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 31.15 | |
| KG-ProverLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 30.74 | |
| KG-ProverLLM Model=GPT 4o, Max Attempts=32025.02 | 30.74 | |
| KG-ProverLLM Model=o1-mini, Max Attempts=32025.02 | 30.74 | |
| RAGLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 28.69 | |
| RAGLLM Model=GPT 4o, Max Attempts=32025.02 | 28.69 | |
| RAGLLM Model=o1-mini, Max Attempts=32025.02 | 28.28 | |
| KG-ProverLLM Model=Deepseek R1, Max Attempts=32025.02 | 28.28 | |
| BaseLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 25 | |
| RAGLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 24.59 | |
| RAGLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 24.59 | |
| BaseLLM Model=o1-mini, Max Attempts=32025.02 | 23.77 | |
| BaseLLM Model=GPT 4o, Max Attempts=32025.02 | 23.36 | |
| BaseLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 22.95 | |
| RAGLLM Model=Deepseek R1, Max Attempts=32025.02 | 22.54 | |
| BaseLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 20.49 | |
| BaseLLM Model=Deepseek R1, Max Attempts=32025.02 | 20.08 |