Question Answering on QA OOD StrQA SciQA
98.3StrQA AccuracyQwen3-8B pass@N (Upper Bound)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen3-8B pass@N (Upper Bound)# Sample=-2025.11 | 98.3 | 99.3 | — | — | |
| Qwen3-14B pass@1# Sample=-2025.11 | 91.9 | 95.6 | — | — | |
| ReProbe, Attn+Logit, Self-anno# Sample=32K, Input Features=Attn+Logit, Annotation Source=Self-anno2025.11 | 88.6 | 97.1 | — | — | |
| ReProbe, Attn+Logit, DeepSeek-anno# Sample=32K, Input Features=Attn+Logit, Annotation Source=DeepSeek-anno2025.11 | 88.6 | 96.9 | — | — | |
| ReProbe, Hidden States, DeepSeek-anno# Sample=32K, Input Features=Hidden States, Annotation Source=DeepSeek-anno2025.11 | 88.6 | 96.9 | — | — | |
| Qwen2.5-Math-PRM-7B# Sample=860K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 88.1 | 96.9 | — | — | |
| ReProbe, Hidden States, Self-anno# Sample=32K, Input Features=Hidden States, Annotation Source=Self-anno2025.11 | 88.1 | 95.8 | — | — | |
| RLHFlow-PRM-Mistral-Data# Sample=273K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 87.8 | 94.1 | — | — | |
| Universal-PRM-Qwen2.5-Math-7B# Sample=690K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 87.8 | 97.1 | — | — | |
| RLHFlow-PRM-Deepseek-Data# Sample=253K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 87.6 | 94.1 | — | — | |
| Math-Shepherd-PRM-7B# Sample=440K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 87.3 | 95.8 | — | — | |
| Qwen2.5-Math-7B-PRM800k# Sample=263K, PRM Category=PRMs 750× to 810× Larger than ReProbes2025.11 | 87.1 | 96.9 | — | — | |
| Qwen3-8B pass@1 (Lower Bound)# Sample=-2025.11 | 86.8 | 92.7 | — | — | |
| Majority Voting# Sample=-2025.11 | 86.6 | 92.5 | — | — | |
| Skywork-PRM-1.5B# Sample=Unk, PRM Category=PRMs 150× Larger than ReProbes2025.11 | 86.6 | 96.3 | — | — | |
| H4-Qwen2.5-PRM-1.5B-0.2# Sample=369K, PRM Category=PRMs 150× Larger than ReProbes2025.11 | 84.6 | 94.7 | — | — | |
| Qwen3-1.7B pass@N (Upper Bound)# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=pass@N, Inference Mode=native thinking mode2025.11 | 74 | 89.2 | 81.6 | 83.4 | |
| Math-Shepherd-PRM-7B# Sample=440K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 54.6 | 47.2 | 50.9 | 62 | |
| ReProbe, Hidden States, GPT-OSS-anno# Sample=10.8K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 51.8 | 45.9 | 48.9 | 60.6 | |
| RLHFlow-PRM-Mistral-Data# Sample=273K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 51.2 | 49.8 | 50.5 | 59 | |
| MaxProb# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 50.9 | 59.7 | 55.3 | 60.6 | |
| Perplexity# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 50.9 | 59.7 | 55.3 | 60.6 | |
| RLHFlow-PRM-Deepseek-Data# Sample=253K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 50 | 45.6 | 47.8 | 58.6 | |
| MaxEntropy# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 49.4 | 60 | 54.7 | 60.5 | |
| Universal-PRM-Qwen2.5-Math-7B# Sample=690K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 48.5 | 57.4 | 53 | 63 | |
| H4-Qwen2.5-PRM-1.5B-0.2# Sample=369K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 45.4 | 47.2 | 46.3 | 57.5 | |
| Qwen2.5-Math-PRM-7B# Sample=860K, Base Model=Qwen3-1.7B, Decoding Strategy=Best-of-N=10, Inference Mode=native thinking mode2025.11 | 44.8 | 51.5 | 48.2 | 59.4 | |
| Qwen3-1.7B pass@1 (Lower Bound)# Sample=–, Base Model=Qwen3-1.7B, Decoding Strategy=pass@1, Inference Mode=native thinking mode2025.11 | 43.2 | 48.3 | 45.8 | 52.8 |