Automated Theorem Proving on MUSTARDSAUCE
34AccuracyKG-Prover
Evaluation Results
| Method | Links | |
|---|---|---|
| KG-ProverLLM Model=o1-mini, Max Attempts=32025.02 | 34 | |
| KG-ProverLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 32.5 | |
| KG-ProverLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 30 | |
| KG-ProverLLM Model=GPT 4o, Max Attempts=32025.02 | 30 | |
| RAGLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 28.8 | |
| RAGLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 28.4 | |
| BaseLLM Model=Claude 3.5 Sonnet, Max Attempts=32025.02 | 28 | |
| BaseLLM Model=GPT 4o, Max Attempts=32025.02 | 28 | |
| RAGLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 28 | |
| RAGLLM Model=GPT 4o, Max Attempts=32025.02 | 28 | |
| KG-ProverLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 27.6 | |
| KG-ProverLLM Model=Deepseek R1, Max Attempts=32025.02 | 27 | |
| RAGLLM Model=o1-mini, Max Attempts=32025.02 | 26.8 | |
| BaseLLM Model=Llama 3.3 70B, Max Attempts=32025.02 | 25.6 | |
| RAGLLM Model=Deepseek R1, Max Attempts=32025.02 | 25 | |
| BaseLLM Model=o1-mini, Max Attempts=32025.02 | 24.8 | |
| BaseLLM Model=Llama 3.1 8B, Max Attempts=32025.02 | 24 | |
| BaseLLM Model=Deepseek R1, Max Attempts=32025.02 | 20 |