Mathematical Reasoning on AIME24 (Accuracy, Average output length)
73.59AccuracyMath-Shepherd-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Math-Shepherd-7BBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 73.59 | 5,820 | |
| OTVBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 73.44 | 5,447 | |
| Qwen2.5-PRM-7BBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 71.77 | 3,720 | |
| Qwen2.5-PRM800K-7BBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 69.9 | 4,891 | |
| OTVBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 67.03 | 4,132 | |
| Qwen2.5-PRM800K-7BBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 66.77 | 2,196 | |
| Qwen2.5-PRM-7BBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 66.46 | 3,173 | |
| Qwen2.5-PRM-7BBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 65.73 | 4,211 | |
| DeepConfBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 64.95 | 9,664 | |
| Qwen2.5-PRM-7BBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 63.8 | 3,040 | |
| VersaPRM-8BBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 63.75 | 3,270 | |
| OTVBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 63.75 | 4,431 | |
| OTVBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 63.39 | 4,427 | |
| OTVBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 63.18 | 4,623 | |
| DeepConfBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 62.86 | 7,044 | |
| Math-Shepherd-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 62.34 | 5,051 | |
| Qwen2.5-PRM800K-7BBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 62.24 | 2,919 | |
| DeepConfBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 61.98 | 6,742 | |
| Math-Shepherd-7BBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 61.61 | 5,679 | |
| VersaPRM-8BBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 60.52 | 2,589 | |
| Qwen2.5-PRM800K-7BBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 60.31 | 3,434 | |
| DeepConfBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 59.9 | 7,150 | |
| VersaPRM-8BBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 59.79 | 5,263 | |
| VersaPRM-8BBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 59.06 | 2,981 | |
| Math-Shepherd-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 58.8 | 4,983 | |
| Math-Shepherd-7BBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 57.45 | 5,989 | |
| Qwen2.5-PRM-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 57.03 | 4,888 | |
| OTVBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 55.95 | 3,397 | |
| Qwen2.5-PRM-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 55.57 | 4,939 | |
| Math-Shepherd-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 55.52 | 4,819 | |
| Qwen2.5-PRM-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 55.52 | 5,525 | |
| Math-Shepherd-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 55.21 | 5,555 | |
| OTVBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 55.05 | 3,436 | |
| Math-Shepherd-7BBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 54.64 | 5,722 | |
| VersaPRM-8BBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 54.48 | 6,560 | |
| Qwen2.5-PRM800K-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 54.17 | 4,722 | |
| DeepConfBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 53.92 | 5,101 | |
| OTVBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 53.54 | 3,211 | |
| Qwen2.5-PRM800K-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 53.28 | 5,585 | |
| VersaPRM-8BBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 53.12 | 5,432 | |
| Qwen2.5-PRM-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 51.98 | 5,660 | |
| DeepConfBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 51.82 | 7,176 | |
| DeepConfBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 50.94 | 3,772 | |
| DeepConfBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 50.52 | 6,382 | |
| Qwen2.5-PRM800K-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 49.48 | 5,260 | |
| VersaPRM-8BBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 49.32 | 5,744 | |
| VersaPRM-8BBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 48.8 | 5,061 | |
| Qwen2.5-PRM800K-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 47.4 | 5,732 |