Step-level Reasoning Verification on Step-level Verification Suite Math, Planning, QA
0.312MATH AccuracyReProbe
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ReProbeNumber of Samples (# Sample)=10.8K, Thinking Mode=native thinking, Input Features=Hidden States, Annotation Source=GPT-OSS-anno, Model Scale=Qwen3-1.7B2025.11 | 0.312 | 0.433 | 0.465 | 0.334 | 0.505 | 0.604 | 0.426 | 0.451 | 0.403 | 0.474 | 0.447 | |
| Qwen2.5-Math-7BNumber of Samples (# Sample)=860K, Thinking Mode=native thinking, Model Scale=Qwen3-1.7B2025.11 | 0.293 | 0.27 | 0.465 | 0.328 | 0.732 | 0.605 | 0.378 | 0.293 | 0.343 | 0.467 | 0.421 | |
| Qwen2.5-Math-7B-PRM800KNumber of Samples (# Sample)=263K, Thinking Mode=native thinking, Model Scale=Qwen3-1.7B2025.11 | 0.217 | 0.392 | 0.505 | 0.27 | 0.716 | 0.587 | 0.391 | 0.336 | 0.371 | 0.46 | 0.427 | |
| Universal-PRM-Qwen2.5-Math-7BNumber of Samples (# Sample)=690K, Thinking Mode=native thinking, Model Scale=Qwen3-1.7B2025.11 | 0.175 | 0.331 | 0.299 | 0.272 | 0.708 | 0.446 | 0.335 | 0.329 | 0.268 | 0.418 | 0.362 | |
| RLHFlow-PRM-Deepseek-DataNumber of Samples (# Sample)=253K, Thinking Mode=native thinking, Model Scale=Qwen3-1.7B2025.11 | 0.063 | 0.267 | 0.228 | 0.169 | 0.498 | 0.322 | 0.449 | 0.368 | 0.186 | 0.361 | 0.296 | |
| Math-Shepherd-PRM-7BNumber of Samples (# Sample)=440K, Thinking Mode=native thinking, Model Scale=Qwen3-1.7B2025.11 | 0.049 | 0.248 | 0.234 | 0.219 | 0.481 | 0.338 | 0.265 | 0.313 | 0.177 | 0.323 | 0.268 | |
| RLHFlow-PRM-Mistral-DataNumber of Samples (# Sample)=273K, Thinking Mode=native thinking, Model Scale=Qwen3-1.7B2025.11 | 0.047 | 0.184 | 0.164 | 0.141 | 0.457 | 0.275 | 0.273 | 0.279 | 0.132 | 0.285 | 0.228 |