Theorem Autoformalization on FormalPhysics
100FVFormalPhysics
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| FormalPhysicsSize=200, Domain=Advanced Physics, sNL (Natural Language Statement)=✓, pNL (Natural Language Proof)=✓, sFL (Formal Language Statement)=✓, pFL (Formal Language Proof)=✓2026.04 | 100 | 6.41 | 6.22 | 73.5 | 72 | 72.5 | |
| Kimina-7BPipeline=Zero-Shot Autoformalisation, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 51.5 | — | — | 11 | 14.5 | 6.5 | |
| Kimina-7BPipeline=Self-Refinement with Error Feedback, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 23 | — | — | 6 | 7.5 | 5 | |
| GPT-5.1Pipeline=Self-Refinement with Error Feedback, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 17 | — | — | 38 | 35 | 42 | |
| GPT-5.1Pipeline=Zero-Shot Autoformalisation, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 14.5 | — | — | 27 | 28 | 33 | |
| DeepSeek-Prover-7BPipeline=Zero-Shot Autoformalisation, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 13 | — | — | 12.5 | 13.5 | 14 | |
| GPT-OSS-20BPipeline=Self-Refinement with Error Feedback, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 7.5 | — | — | 14.5 | 10.5 | 16.5 | |
| GPT-OSS-20BPipeline=Zero-Shot Autoformalisation, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 4.5 | — | — | 15.5 | 12.5 | 17.5 | |
| DeepSeek-Prover-7BPipeline=Self-Refinement with Error Feedback, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 4.5 | — | — | 26.5 | 11.5 | 17 | |
| Qwen2.5-Coder-7BPipeline=Zero-Shot Autoformalisation, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 1 | — | — | 8 | 9 | 12.5 | |
| Qwen2.5-Coder-7BPipeline=Self-Refinement with Error Feedback, Evaluator=Qwen2.5-Coder-7B-Instruct2026.04 | 1 | — | — | 11.5 | 7 | 10.5 |