Formal Mathematical Proving on OptBench
55.37Basic Pass@32OptProver (EI + PW-UAPO)
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OptProver (EI + PW-UAPO)Proving Strategy=Step-level, Parameter Scale=7B, Expert Iteration=EI, Optimization Objective=PW-UAPO2026.04 | 55.37 | 40.5 | — | 45.19 | 62.22 | — | 28.47 | 48.61 | — | 37.75 | 55.25 | — | |
| OptProver (EI + UAPO)Proving Strategy=Step-level, Parameter Scale=7B, Expert Iteration=EI, Optimization Objective=UAPO2026.04 | 53.72 | 36.36 | — | 45.93 | 60 | — | 27.08 | 47.22 | — | 36.25 | 53.5 | — | |
| OptProver (EI + PW-DPO)Proving Strategy=Step-level, Parameter Scale=7B, Expert Iteration=EI, Optimization Objective=PW-DPO2026.04 | 53.72 | 37.19 | — | 45.19 | 62.96 | — | 27.08 | 46.53 | — | 36.25 | 54.25 | — | |
| OptProver (EI + DPO)Proving Strategy=Step-level, Parameter Scale=7B, Expert Iteration=EI, Optimization Objective=DPO2026.04 | 52.07 | 34.71 | — | 39.26 | 62.22 | — | 20.14 | 42.36 | — | 31 | 52 | — | |
| OptProver (Only EI)Proving Strategy=Step-level, Parameter Scale=7B, Expert Iteration=Only EI2026.04 | 49.59 | 32.23 | — | 39.26 | 60 | — | 15.97 | 40.97 | — | 28.75 | 50 | — | |
| BFS-Prover-V2-7BProving Strategy=Step-level, Parameter Scale=7B, Search Algorithm=Best-first search2026.04 | 36.36 | 14.05 | — | 20 | 44.44 | — | 6.25 | 16.67 | — | 13.25 | 32 | — | |
| DeepSeek-Prover-V2-7BProving Strategy=Whole-proof, Parameter Scale=7B2026.04 | 14.88 | — | 19.01 | — | 23.7 | 31.85 | — | 4.86 | 6.25 | — | 14.25 | 18.75 | |
| Goedel-Prover-V2-8BProving Strategy=Whole-proof, Parameter Scale=8B2026.04 | 14.05 | — | 19.01 | — | 19.26 | 32.59 | — | 4.86 | 8.33 | — | 12.5 | 19.75 | |
| Kimina-Prover-7BProving Strategy=Whole-proof, Parameter Scale=7B2026.04 | 4.96 | — | 10.74 | — | 10.37 | 16.3 | — | 2.08 | 5.56 | — | 5.75 | 10.75 |