Mathematical Reasoning on Pooled 5-benchmark set
54.44AccuracyLEAD
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LEADBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 54.44 | 5,620 | 0.54 | |
| BaseBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 51.85 | 9,213 | — | |
| DRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 51.34 | 4,635 | 0.4 | |
| GDPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 51.27 | 3,252 | 0.53 | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 50.03 | 3,007 | 0.32 | |
| ShorterBetterBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 44.59 | 2,713 | 0.69 |