Step-wise Confidence Attribution on GSM8K
0.7892AUROCGIBS
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GIBSLLM=Phi4-reasoning2026.05 | 0.7892 | 0.8172 | 81.17 | 0.3354 | |
| NLI-MaxLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.745 | 0.6982 | 64.09 | 0.1456 | |
| GIBSLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.7289 | 0.6712 | 65.32 | 0.2867 | |
| NLI-MaxLLM=Llama3.1-8b2026.05 | 0.7096 | 0.789 | 59.08 | 0.1162 | |
| GIBSLLM=Llama3.1-8b2026.05 | 0.691 | 0.7004 | 62.92 | 0.2293 | |
| NLI-MeanLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.6762 | 0.6508 | 63.18 | 0.4189 | |
| Cos-MeanLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.6633 | 0.6933 | 57.48 | 0.3323 | |
| NLI-MaxLLM=Phi4-reasoning2026.05 | 0.66 | 0.8141 | 74.46 | 0.2009 | |
| Cos-MeanLLM=Llama3.1-8b2026.05 | 0.6078 | 0.6211 | 61.75 | 0.33 | |
| Cos-MeanLLM=Phi4-reasoning2026.05 | 0.5959 | 0.7703 | 71.99 | 0.1556 | |
| NLI-MeanLLM=Phi4-reasoning2026.05 | 0.5738 | 0.7704 | 71.86 | 0.5527 | |
| NLI-MeanLLM=Llama3.1-8b2026.05 | 0.5524 | 0.6103 | 56.65 | 0.4376 | |
| Cos-MaxLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.5269 | 0.5197 | 56.76 | 0.3799 | |
| P(true)LLM=Phi4-reasoning2026.05 | 0.5251 | 0.6956 | 69.34 | 0.6851 | |
| EntropyLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.5203 | 0.555 | 52.35 | 0.2583 | |
| P(true)LLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.5159 | 0.582 | 58.4 | 0.5802 | |
| SL(norm)LLM=Llama3.1-8b2026.05 | 0.479 | 0.524 | 55.13 | 0.2282 | |
| EntropyLLM=Phi4-reasoning2026.05 | 0.4623 | 0.6795 | 66.64 | 0.1648 | |
| Cos-MaxLLM=Llama3.1-8b2026.05 | 0.4537 | 0.5152 | 55.13 | 0.378 | |
| EntropyLLM=Llama3.1-8b2026.05 | 0.4105 | 0.4962 | 51.41 | 0.2518 | |
| P(true)LLM=Llama3.1-8b2026.05 | 0.4016 | 0.4711 | 52.83 | 0.5504 | |
| LECOLLM=Llama3.1-8b2026.05 | 0.3862 | 0.4586 | 53.19 | 0.3783 | |
| SL(norm)LLM=Phi4-reasoning2026.05 | 0.3851 | 0.5947 | 67.99 | 0.2394 | |
| SL(norm)LLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.37 | 0.5308 | 52.62 | 0.123 | |
| Cos-MaxLLM=Phi4-reasoning2026.05 | 0.3494 | 0.5861 | 70.1 | 0.1997 | |
| LECOLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.3202 | 0.411 | 48.83 | 0.3395 | |
| LECOLLM=Phi4-reasoning2026.05 | 0.2885 | 0.5735 | 63.44 | 0.4021 |