Mathematical Reasoning Verification on MATH 500
48.4Accuracy (MATH 500)LMUnit-qwen2.5-72B
Evaluation Results
| Method | Links | |
|---|---|---|
| LMUnit-qwen2.5-72BType=Trained verifiable reward model, Backbone=Qwen2.5-72B2025.08 | 48.4 | |
| Skywork-Reward-V2-Llama-3.1-8BType=Trained verifiable reward model, Backbone=Llama-3.1-8B2025.08 | 48.4 | |
| PiCSARzero-shot=true, training-free=true2025.08 | 46.53 |