Mathematical Reasoning on Math ID GSM8k ProofNet
99.2GSM8k AccuracyQwen3-8B pass@N (Upper Bound)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen3-8B pass@N (Upper Bound)# Sample=-2025.11 | 99.2 | 99.3 | 97.8 | — | |
| Qwen2.5-Math-PRM-7B# Sample=860K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 97.8 | 93.7 | 76 | — | |
| ReProbe, Attn+Logit, DeepSeek-anno# Sample=32K, Input Features=Attn+Logit, Annotation Source=DeepSeek-anno2025.11 | 97.8 | 92.7 | 76.5 | — | |
| Qwen3-14B pass@1# Sample=-2025.11 | 97.6 | 93.4 | 76 | — | |
| Majority Voting# Sample=-2025.11 | 97.6 | — | — | — | |
| Skywork-PRM-1.5B# Sample=Unk, PRM Category=PRMs 150× Larger than ReProbes2025.11 | 97.6 | 94.4 | 76.5 | — | |
| Universal-PRM-Qwen2.5-Math-7B# Sample=690K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 97.5 | 95.7 | 76 | — | |
| ReProbe, Attn+Logit, Self-anno# Sample=32K, Input Features=Attn+Logit, Annotation Source=Self-anno2025.11 | 97.5 | 94.4 | 73.6 | — | |
| Qwen2.5-Math-7B-PRM800k# Sample=263K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 97.3 | 92.7 | 74.4 | — | |
| RLHFlow-PRM-Deepseek-Data# Sample=253K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 96.4 | 92.7 | 71.7 | — | |
| RLHFlow-PRM-Mistral-Data# Sample=273K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 96.3 | 93.7 | 71.7 | — | |
| ReProbe, Hidden States, Self-anno# Sample=32K, Input Features=Hidden States, Annotation Source=Self-anno2025.11 | 96.1 | 94.4 | 78.5 | — | |
| ReProbe, Hidden States, DeepSeek-anno# Sample=32K, Input Features=Hidden States, Annotation Source=DeepSeek-anno2025.11 | 96.1 | 92.7 | 74.1 | — | |
| Qwen3-8B pass@1 (Lower Bound)# Sample=-2025.11 | 95.6 | 92.4 | 74.1 | — | |
| Math-Shepherd-PRM-7B# Sample=440K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 95.5 | 93 | 72.8 | — | |
| H4-Qwen2.5-PRM-1.5B-0.2# Sample=369K, PRM Category=PRMs 150× Larger than ReProbes2025.11 | 95.1 | 91.7 | 71.4 | — | |
| Qwen3-1.7B pass@N (Upper Bound)# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=pass@N, Inference Mode=native thinking mode2025.11 | 93.3 | 83.1 | 77.4 | 84.6 | |
| Qwen2.5-Math-PRM-7B# Sample=860K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 85.4 | 62.6 | 52.9 | 67 | |
| Universal-PRM-Qwen2.5-Math-7B# Sample=690K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 85 | 67.3 | 57 | 69.8 | |
| ReProbe, Hidden States, GPT-OSS-anno# Sample=10.8K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 85 | 62.2 | 57.9 | 68.4 | |
| Math-Shepherd-PRM-7B# Sample=440K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 84.6 | 67.3 | 56.2 | 69.4 | |
| MaxProb# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 84 | 55.6 | 52.9 | 64.2 | |
| Perplexity# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 84 | 55.6 | 52.9 | 64.2 | |
| MaxEntropy# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 83.6 | 62.2 | 47.1 | 64.3 | |
| RLHFlow-PRM-Deepseek-Data# Sample=253K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 82.8 | 64.9 | 49.6 | 65.8 | |
| RLHFlow-PRM-Mistral-Data# Sample=273K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 82.4 | 62.6 | 48.8 | 64.6 | |
| H4-Qwen2.5-PRM-1.5B-0.2# Sample=369K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 78.3 | 63.7 | 52.9 | 65 | |
| Qwen3-1.7B pass@1 (Lower Bound)# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=pass@1, Inference Mode=native thinking mode2025.11 | 70.1 | 54.6 | 47.7 | 57.5 |