Reasoning on MATH (AUROC/FPR95)
0.8495AUROCTRACED
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TRACEDModel=Qwen3-4B-Thinking-25072026.03 | 0.8495 | 0.725 | 0.8422 | |
| SAPLMAModel=Qwen3-4B-Thinking-25072026.03 | 0.8438 | 0.7572 | 0.8406 | |
| KV-CoE-CModel=Qwen2-7B-Instruct2026.01 | 0.8412 | 0.4482 | — | |
| LR ProbeModel=Qwen3-4B-Thinking-25072026.03 | 0.7906 | 0.75 | 0.8344 | |
| KV-CoE-RModel=Qwen2-7B-Instruct2026.01 | 0.7692 | 0.4983 | — | |
| CoE-CModel=Qwen2-7B-Instruct2026.01 | 0.7668 | 0.6448 | — | |
| CoE-RModel=Qwen2-7B-Instruct2026.01 | 0.7575 | 0.6595 | — | |
| CoEModel=Qwen3-4B-Thinking-25072026.03 | 0.75 | 0.7542 | 0.8411 | |
| TRACEDModel=DeepSeek-R1-Llama-8B2026.03 | 0.7489 | 0.75 | 0.6549 | |
| LR ProbeModel=DeepSeek-R1-Llama-8B2026.03 | 0.7471 | 0.8625 | 0.6541 | |
| CoE-CModel=Llama3-8B2026.01 | 0.7308 | 0.796 | — | |
| TRACEDModel=Qwen2.5-7B-Instruct2026.03 | 0.7305 | 0.828 | 0.7859 | |
| LR ProbeModel=Qwen2.5-7B-Instruct2026.03 | 0.7258 | 0.8378 | 0.7349 | |
| CoE-RModel=Llama3-8B2026.01 | 0.7254 | 0.7561 | — | |
| SAPLMAModel=DeepSeek-R1-Llama-8B2026.03 | 0.7161 | 0.8 | 0.6363 | |
| SAPLMAModel=Qwen2.5-7B-Instruct2026.03 | 0.7003 | 0.8384 | 0.7538 | |
| CoT-KineticsModel=DeepSeek-R1-Llama-8B2026.03 | 0.6755 | 0.7933 | 0.6271 | |
| KV-CoE-RModel=Llama-3.1-8B-Instruct2026.01 | 0.6436 | 0.6382 | — | |
| KV-CoE-CModel=Llama-3.1-8B-Instruct2026.01 | 0.6413 | 0.6742 | — | |
| TRACEDModel=Llama-3.1-8B-Instruct2026.03 | 0.6363 | 0.8 | 0.6265 | |
| LR ProbeModel=Llama-3.1-8B-Instruct2026.03 | 0.6362 | 0.8667 | 0.6025 | |
| MSPModel=DeepSeek-R1-Llama-8B2026.03 | 0.6276 | 0.8125 | 0.6203 | |
| EntropyModel=Llama-3.1-8B-Instruct2026.01 | 0.6274 | 0.8414 | — | |
| SAPLMAModel=Llama-3.1-8B-Instruct2026.03 | 0.6195 | 0.8333 | 0.6046 | |
| CoEModel=DeepSeek-R1-Llama-8B2026.03 | 0.6156 | 0.8812 | 0.6332 | |
| PerplexityModel=Qwen3-4B-Thinking-25072026.03 | 0.6094 | 0.765 | 0.5602 | |
| PPLModel=Llama-3.1-8B-Instruct2026.01 | 0.6082 | 0.8642 | — | |
| EntropyModel=DeepSeek-R1-Llama-8B2026.03 | 0.6041 | 0.8875 | 0.5921 | |
| MaxProbModel=Llama-3.1-8B-Instruct2026.01 | 0.5916 | 0.8796 | — | |
| EntropyModel=Llama-3.1-8B-Instruct2026.03 | 0.5822 | 0.8333 | 0.5795 | |
| PerplexityModel=DeepSeek-R1-Llama-8B2026.03 | 0.5808 | 0.8688 | 0.5628 | |
| MSPModel=Llama-3.1-8B-Instruct2026.03 | 0.5777 | 0.8167 | 0.6196 | |
| CoEModel=Qwen2.5-7B-Instruct2026.03 | 0.5586 | 0.8863 | 0.6308 | |
| PerplexityModel=Qwen2.5-7B-Instruct2026.03 | 0.5567 | 0.8317 | 0.6053 | |
| CoT-KineticsModel=Qwen2.5-7B-Instruct2026.03 | 0.5442 | 0.8386 | 0.5223 | |
| CoEModel=Llama-3.1-8B-Instruct2026.03 | 0.5373 | 0.8333 | 0.5193 | |
| PerplexityModel=Llama-3.1-8B-Instruct2026.03 | 0.5339 | 0.85 | 0.5197 | |
| EntropyModel=Qwen3-4B-Thinking-25072026.03 | 0.5312 | 0.737 | 0.5285 | |
| MSPModel=Qwen3-4B-Thinking-25072026.03 | 0.4844 | 0.75 | 0.5592 | |
| MSPModel=Qwen2.5-7B-Instruct2026.03 | 0.4828 | 0.8833 | 0.5201 | |
| EntropyModel=Qwen2.5-7B-Instruct2026.03 | 0.4817 | 0.8842 | 0.52 | |
| CoT-KineticsModel=Llama-3.1-8B-Instruct2026.03 | 0.4367 | 0.8467 | 0.4991 | |
| CoT-KineticsModel=Qwen3-4B-Thinking-25072026.03 | 0.4219 | 0.85 | 0.5 | |
| EntropyModel=Qwen2-7B-Instruct2026.01 | 0.1619 | 0.9942 | — | |
| PPLModel=Qwen2-7B-Instruct2026.01 | 0.1243 | 0.995 | — | |
| MaxProbModel=Qwen2-7B-Instruct2026.01 | 0.124 | 0.9934 | — | |
| H4-Qwen2.5-PRM-1.5B-0.2# Sample=369K, Category=PRMs 150x Larger than ReProbes2025.11 | — | — | 0.174 | |
| Math-Shepherd-PRM-7B# Sample=440K, Category=PRMs 750x to 810x Larger than ReProbes2025.11 | — | — | 0.248 | |
| MaxEntropy# Sample=-, Category=Unsupervised Uncertainty Quantification (UQ)2025.11 | — | — | 0.112 | |
| MaxProb# Sample=-, Category=Unsupervised Uncertainty Quantification (UQ)2025.11 | — | — | 0.127 | |
| Perplexity# Sample=-, Category=Unsupervised Uncertainty Quantification (UQ)2025.11 | — | — | 0.117 | |
| Qwen2.5-Math-7B# Sample=860K, Category=PRMs 750x to 810x Larger than ReProbes2025.11 | — | — | 0.427 | |
| Qwen2.5-Math-7B-PRM800k# Sample=265K, Category=PRMs 750x to 810x Larger than ReProbes2025.11 | — | — | 0.474 | |
| Random# Sample=-, Category=Unsupervised Uncertainty Quantification (UQ)2025.11 | — | — | 0.106 | |
| ReProbe, Attn+Logit, Qwen3-8B-anno# Sample=32K, Category=Reasoning Probes (ReProbes)2025.11 | — | — | 0.404 | |
| RLHFlow-PRM-Deepseek-8B# Sample=253K, Category=PRMs 750x to 810x Larger than ReProbes2025.11 | — | — | 0.2 | |
| RLHFlow-PRM-Mistral-8B# Sample=273K, Category=PRMs 750x to 810x Larger than ReProbes2025.11 | — | — | 0.141 | |
| Skywork-PRM-1.5B# Sample=Unk, Category=PRMs 150x Larger than ReProbes2025.11 | — | — | 0.219 | |
| Universal-PRM-Qwen2.5-Math-7B# Sample=690K, Category=PRMs 750x to 810x Larger than ReProbes2025.11 | — | — | 0.485 |