Statement Autoformalisation on FormalPhysics corpus
100FV ScoreFormalScience
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| FormalScienceInference Pipeline=(4) FormalScience (ours), Evaluator LLM=GPT-4.1-mini, Models=GPT-5.1 / Claude-4.52026.04 | 100 | 73.5 | 72 | 72.5 | |
| Qwen3-Sonnet-14BInference Pipeline=(3) Agentic Code Generation Pipeline, Evaluator LLM=GPT-4.1-mini2026.04 | 52 | 1 | 10.5 | 6.5 | |
| Kimina-7BInference Pipeline=(1) Zero-Shot Autoformalisation, Evaluator LLM=GPT-4.1-mini2026.04 | 51.5 | 6.5 | 10.5 | 9.5 | |
| GPT-OSS-20BInference Pipeline=(3) Agentic Code Generation Pipeline, Evaluator LLM=GPT-4.1-mini2026.04 | 31 | 73 | 72.5 | 73 | |
| Kimina-7BInference Pipeline=(2) Self-Refinement with Error Feedback, Evaluator LLM=GPT-4.1-mini2026.04 | 23 | 6.5 | 9.5 | 8 | |
| GPT-5.1Inference Pipeline=(2) Self-Refinement with Error Feedback, Evaluator LLM=GPT-4.1-mini2026.04 | 17 | 82.5 | 82 | 82 | |
| GPT-5.1Inference Pipeline=(1) Zero-Shot Autoformalisation, Evaluator LLM=GPT-4.1-mini2026.04 | 14.5 | 79.5 | 76.5 | 77 | |
| DeepSeek-Prover-7BInference Pipeline=(1) Zero-Shot Autoformalisation, Evaluator LLM=GPT-4.1-mini2026.04 | 13 | 23 | 27.5 | 24 | |
| GPT-OSS-20BInference Pipeline=(2) Self-Refinement with Error Feedback, Evaluator LLM=GPT-4.1-mini2026.04 | 7.5 | 70.5 | 77 | 79 | |
| Qwen3-Coder-30BInference Pipeline=(3) Agentic Code Generation Pipeline, Evaluator LLM=GPT-4.1-mini2026.04 | 5.5 | 49.5 | 59 | 48 | |
| GPT-OSS-20BInference Pipeline=(1) Zero-Shot Autoformalisation, Evaluator LLM=GPT-4.1-mini2026.04 | 4.5 | 68.5 | 73 | 72.5 | |
| DeepSeek-Prover-7BInference Pipeline=(2) Self-Refinement with Error Feedback, Evaluator LLM=GPT-4.1-mini2026.04 | 4.5 | 17 | 23 | 23 | |
| Qwen2.5-Coder-7BInference Pipeline=(1) Zero-Shot Autoformalisation, Evaluator LLM=GPT-4.1-mini2026.04 | 1 | 15 | 24 | 20.5 | |
| Qwen2.5-Coder-7BInference Pipeline=(2) Self-Refinement with Error Feedback, Evaluator LLM=GPT-4.1-mini2026.04 | 1 | 16.5 | 23 | 19.5 |