Mathematical Proof Reward Modeling on Proof-RM IMO CMO USAMO 2024-2025 (test)
82.4AccuracyProofRM-32B
Evaluation Results
| Method | Links | |
|---|---|---|
| ProofRM-32BParameters=32B2026.02 | 82.4 | |
| ProofRM-14BParameters=14B2026.02 | 79 | |
| ProofRM-8BParameters=8B2026.02 | 76.8 | |
| GPT-5-mini-2025-08-07Model Version=mini-2025-08-072026.02 | 74.6 | |
| Gemini-2.5-FlashModel Version=2.5-Flash2026.02 | 74.5 | |
| Distill-Qwen3-32BBase Model=Qwen3, Parameters=32B, Type=Distilled2026.02 | 72.5 | |
| Distill-Qwen3-14BBase Model=Qwen3, Parameters=14B, Type=Distilled2026.02 | 71.9 | |
| Distill-Qwen3-8BBase Model=Qwen3, Parameters=8B, Type=Distilled2026.02 | 71.6 | |
| Deepseek V3.1Model Version=V3.12026.02 | 69.6 |