Mathematical Reasoning on AIME25 (Accuracy, Average output length)
66.51AccuracyMath-Shepherd-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Math-Shepherd-7BBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 66.51 | 5,824 | |
| OTVBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 53.91 | 5,416 | |
| Qwen2.5-PRM-7BBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 53.33 | 3,948 | |
| OTVBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 50.68 | 2,926 | |
| Qwen2.5-PRM800K-7BBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 49.53 | 6,409 | |
| OTVBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 49.11 | 6,542 | |
| OTVBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 49.02 | 3,170 | |
| OTVBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 48.23 | 2,577 | |
| Qwen2.5-PRM800K-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 47.81 | 4,426 | |
| Math-Shepherd-7BBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 47.45 | 6,218 | |
| Qwen2.5-PRM-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 47.24 | 4,481 | |
| OTVBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 47.08 | 4,079 | |
| OTVBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 46.98 | 2,991 | |
| VersaPRM-8BBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 46.77 | 6,113 | |
| OTVBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 46.46 | 3,225 | |
| Math-Shepherd-7BBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 46.25 | 6,116 | |
| Qwen2.5-PRM-7BBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 45.83 | 6,588 | |
| Qwen2.5-PRM-7BBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 45.73 | 6,416 | |
| Math-Shepherd-7BBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 45.36 | 6,118 | |
| Qwen2.5-PRM800K-7BBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 45.21 | 6,619 | |
| Qwen2.5-PRM800K-7BBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 45.1 | 6,202 | |
| Qwen2.5-PRM-7BBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 44.95 | 6,304 | |
| DeepConfBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 43.54 | 6,318 | |
| VersaPRM-8BBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 43.44 | 6,334 | |
| VersaPRM-8BBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 43.28 | 6,132 | |
| DeepConfBackbone=QWEN3-4B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 43.07 | 9,322 | |
| Qwen2.5-PRM800K-7BBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 42.86 | 6,445 | |
| Math-Shepherd-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 42.76 | 4,919 | |
| Math-Shepherd-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 42.4 | 4,570 | |
| Qwen2.5-PRM800K-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 41.2 | 4,689 | |
| DeepConfBackbone=QWEN3-4B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 40.89 | 7,084 | |
| Qwen2.5-PRM800K-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 40.89 | 4,523 | |
| DeepConfBackbone=QWEN3-4B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 40.78 | 7,260 | |
| Qwen2.5-PRM800K-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 40 | 4,528 | |
| VersaPRM-8BBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 39.48 | 4,796 | |
| Math-Shepherd-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 39.22 | 4,475 | |
| DeepConfBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 38.91 | 3,957 | |
| DeepConfBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 37.76 | 4,353 | |
| VersaPRM-8BBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 37.66 | 5,005 | |
| VersaPRM-8BBackbone=QWEN3-4B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 37.24 | 6,438 | |
| DeepConfBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 37.08 | 4,398 | |
| DeepConfBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 36.77 | 4,449 | |
| VersaPRM-8BBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 36.61 | 5,046 | |
| Math-Shepherd-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 36.46 | 5,063 | |
| Qwen2.5-PRM-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=HALVE@300, N (Sample Count)=1282026.03 | 36.09 | 4,967 | |
| Qwen2.5-PRM-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=STOP@600, N (Sample Count)=1282026.03 | 35.31 | 5,068 | |
| Qwen2.5-PRM-7BBackbone=DAPO-QWEN-32B, Decoding Strategy=DROP@10, N (Sample Count)=1282026.03 | 33.65 | 5,351 | |
| VersaPRM-8BBackbone=DAPO-QWEN-32B, Decoding Strategy=BEST-OF-N, N (Sample Count)=1282026.03 | 31.04 | 4,447 |