Math Proof Verification on ProofBench
69.4AccuracyProofRM-32B
Evaluation Results
| Method | Links | |
|---|---|---|
| ProofRM-32BModel Size=32B2026.02 | 69.4 | |
| ProofRM-14BModel Size=14B2026.02 | 68.5 | |
| ProofRM-8BModel Size=8B2026.02 | 67.5 | |
| Gemini 3 ProTraining/Evaluation Protocol=proprietary2026.04 | 66.7 | |
| QED-Nano (+ RSA test-time scaffold)Model Size=4B, Training/Evaluation Protocol=RSA test-time scaffold2026.04 | 62.6 | |
| DeepSeek-Math-V2Model Size=685B2026.04 | 60.6 | |
| Distill-Qwen3-14BModel Size=14B2026.02 | 59.8 | |
| Distill-Qwen3-8BModel Size=8B2026.02 | 59.3 | |
| Distill-Qwen3-32BModel Size=32B2026.02 | 56.9 | |
| GPT-OSS-120BModel Size=120B2026.04 | 47.5 | |
| QED-NanoModel Size=4B2026.04 | 44.9 | |
| GPT-OSS-20BModel Size=20B2026.04 | 38.4 | |
| Qwen3-235B-A22B-Thinking-2507Model Size=235B-A22B2026.04 | 33.7 | |
| QED-Nano (SFT initialization only)Model Size=4B, Training/Evaluation Protocol=SFT initialization only2026.04 | 33.3 | |
| Nomos-12026.04 | 28.3 | |
| Qwen3-30B-A3B-Thinking-2507Model Size=30B-A3B2026.04 | 26.1 | |
| Qwen3-4B-Thinking-2507Model Size=4B2026.04 | 19.5 |