Multi-step theorem prediction on FormalGeo7K (test)
89.29Total AccuracyPri-TPG
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Pri-TPGMethod category=Neural-symbolic solvers (training-free), Backbone model=GPT-5.2, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 89.29 | 99.16 | 96.28 | 87.92 | 77.07 | 66.13 | 30 | |
| FGeo-HyperGNetMethod category=Neural-symbolic solvers (training-based), Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 88.36 | 96.24 | 91.76 | 87.59 | 82.17 | 56.45 | 56.67 | |
| Pri-TPGMethod category=Neural-symbolic solvers (training-free), Backbone model=GPT-5 mini, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 84.42 | 98.54 | 93.09 | 78.2 | 64.33 | 58.06 | 23.33 | |
| FGeo-TPMethod category=Neural-symbolic solvers (training-based), Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 80.86 | 96.43 | 85.44 | 76.12 | 62.26 | 48.88 | 29.55 | |
| FGeo-DRLMethod category=Neural-symbolic solvers (training-based), Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 80.85 | 97.61 | 91.88 | 70.82 | 57.55 | 36.17 | 27.59 | |
| Claude 4.5 SonnetMethod category=LLM-only (direct solving), Backbone model=Claude 4.5 Sonnet, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 75.79 | 84.55 | 73.94 | 76.32 | 67.52 | 64.52 | 48.33 | |
| GPT-5.2Method category=LLM-only (direct solving), Backbone model=GPT-5.2, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 73.14 | 80.38 | 73.4 | 74.81 | 63.06 | 59.68 | 46.67 | |
| Doubao seed 1.8Method category=LLM-only (direct solving), Backbone model=Doubao seed 1.8, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 69.14 | 74.11 | 69.15 | 71.43 | 64.33 | 50 | 51.67 | |
| Qwen3-VLMethod category=LLM-only (direct solving), Backbone model=Qwen3-VL, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 65.93 | 74.53 | 65.43 | 72.18 | 50.96 | 41.94 | 36.67 | |
| GPT-5 miniMethod category=LLM-only (direct solving), Backbone model=GPT-5 mini, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 64.79 | 74.11 | 63.3 | 64.66 | 53.5 | 53.23 | 41.46 | |
| NGSMethod category=Neural-symbolic solvers (training-based), Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 62.6 | 62.22 | 64.97 | 72.79 | 57.47 | 56.41 | 36.59 | |
| DualGeoSolverMethod category=Neural-symbolic solvers (training-based), Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 62.11 | 62.96 | 67.8 | 65.44 | 60.92 | 53.85 | 34.15 | |
| Inter-GPSMethod category=Neural-symbolic solvers (training-based), Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 60.5 | 76.2 | 63.3 | 60.9 | 39.49 | 17.74 | 15 | |
| DeepSeek v3.2Method category=LLM-only (direct solving), Backbone model=DeepSeek v3.2, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 57.79 | 69.94 | 52.93 | 58.65 | 49.04 | 37.1 | 31.67 | |
| ForwardSearchMethod category=Neural-symbolic solvers (training-free), Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 39.71 | 58.47 | 41.01 | 34.16 | 16.4 | 5.45 | 4.79 | |
| BackwardSearchMethod category=Neural-symbolic solvers (training-free), Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 35.44 | 66.43 | 34.98 | 11.78 | 6.56 | 6.09 | 1.03 | |
| Vanilla ICLMethod category=Neural-symbolic solvers (training-free), Backbone model=GPT-5 mini, Per-problem timeout=600s, Input representation=ground-truth parsed formal inputs2026.03 | 26.29 | 52.19 | 23.67 | 7.89 | 5.1 | 0 | 0 |