Mathematical Verification Reward Modeling on Lean 4 (hold-out set)
0.312Logic MSEValue-Head RM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Value-Head RMModel Strategy=Linear value-head2026.05 | 0.312 | 0.345 | 0.782 | 81.2 | |
| Leibniz-1.5BModel Strategy=Expected Value Alignment (EVA)2026.05 | 0.334 | 0.368 | 0.824 | 84.6 | |
| Standard SFT Generative RMModel Strategy=Fine-tuned using only LSFT2026.05 | 0.485 | 0.521 | 0.725 | 78.5 | |
| GPT-4oModel Strategy=Zero-Shot2026.05 | 0.842 | 1.154 | 0.612 | 72.4 | |
| Qwen2.5-1.5BModel Strategy=Zero-Shot2026.05 | 1.25 | 1.68 | 0.45 | 58.1 |