Reasoning Quality Assessment on TheoremQA
0.873AUROCTRACED
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TRACEDModel=DeepSeek-R1-Llama-8B2026.03 | 0.873 | 0.7094 | 0.7625 | |
| SAPLMAModel=DeepSeek-R1-Llama-8B2026.03 | 0.8518 | 0.6617 | 0.8421 | |
| LR ProbeModel=DeepSeek-R1-Llama-8B2026.03 | 0.8435 | 0.6497 | 0.8158 | |
| SAPLMAModel=Qwen3-4B-Thinking-25072026.03 | 0.7909 | 0.7847 | 0.6327 | |
| TRACEDModel=Qwen2.5-7B-Instruct2026.03 | 0.7752 | 0.7314 | 0.6125 | |
| TRACEDModel=Qwen3-4B-Thinking-25072026.03 | 0.7638 | 0.6333 | 0.625 | |
| LR ProbeModel=Qwen2.5-7B-Instruct2026.03 | 0.7583 | 0.6957 | 0.875 | |
| LR ProbeModel=Qwen3-4B-Thinking-25072026.03 | 0.7364 | 0.7493 | 0.6455 | |
| SAPLMAModel=Qwen2.5-7B-Instruct2026.03 | 0.6958 | 0.7262 | 0.8375 | |
| CoT-KineticsModel=DeepSeek-R1-Llama-8B2026.03 | 0.6738 | 0.5951 | 0.8133 | |
| CoEModel=Qwen2.5-7B-Instruct2026.03 | 0.6625 | 0.6242 | 0.7015 | |
| TRACEDModel=Llama-3.1-8B-Instruct2026.03 | 0.655 | 0.675 | 0.725 | |
| CoEModel=Qwen3-4B-Thinking-25072026.03 | 0.6545 | 0.5495 | 0.6954 | |
| SAPLMAModel=Llama-3.1-8B-Instruct2026.03 | 0.6471 | 0.6551 | 0.7308 | |
| CoT-KineticsModel=Qwen3-4B-Thinking-25072026.03 | 0.6455 | 0.6143 | 0.8394 | |
| LR ProbeModel=Llama-3.1-8B-Instruct2026.03 | 0.6435 | 0.6065 | 0.8077 | |
| PerplexityModel=Qwen3-4B-Thinking-25072026.03 | 0.6273 | 0.5554 | 0.6364 | |
| CoEModel=DeepSeek-R1-Llama-8B2026.03 | 0.6053 | 0.6402 | 0.8474 | |
| CoT-KineticsModel=Llama-3.1-8B-Instruct2026.03 | 0.5754 | 0.5 | 0.85 | |
| MSPModel=DeepSeek-R1-Llama-8B2026.03 | 0.527 | 0.4951 | 0.8737 | |
| PerplexityModel=DeepSeek-R1-Llama-8B2026.03 | 0.5229 | 0.5552 | 0.8211 | |
| PerplexityModel=Llama-3.1-8B-Instruct2026.03 | 0.5133 | 0.5468 | 0.8661 | |
| EntropyModel=DeepSeek-R1-Llama-8B2026.03 | 0.5111 | 0.4951 | 0.8737 | |
| CoEModel=Llama-3.1-8B-Instruct2026.03 | 0.4808 | 0.4835 | 0.8503 | |
| EntropyModel=Llama-3.1-8B-Instruct2026.03 | 0.4719 | 0.5318 | 0.8615 | |
| MSPModel=Llama-3.1-8B-Instruct2026.03 | 0.463 | 0.523 | 0.8589 | |
| MSPModel=Qwen2.5-7B-Instruct2026.03 | 0.4583 | 0.4859 | 0.8375 | |
| EntropyModel=Qwen2.5-7B-Instruct2026.03 | 0.4458 | 0.4725 | 0.8375 | |
| CoT-KineticsModel=Qwen2.5-7B-Instruct2026.03 | 0.4417 | 0.5 | 0.8906 | |
| PerplexityModel=Qwen2.5-7B-Instruct2026.03 | 0.3958 | 0.5279 | 0.8375 | |
| MSPModel=Qwen3-4B-Thinking-25072026.03 | 0.3273 | 0.4622 | 0.8091 | |
| EntropyModel=Qwen3-4B-Thinking-25072026.03 | 0.3273 | 0.4622 | 0.8091 |