Step-wise Confidence Attribution on MoreHopQA
0.8084AUROCGIBS
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GIBSLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.8084 | 0.8357 | 70.51 | 0.1832 | |
| NLI-MaxLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.6663 | 0.7766 | 64.57 | 0.1107 | |
| GIBSLLM=Phi4-reasoning2026.05 | 0.6619 | 0.6866 | 70.53 | 0.356 | |
| GIBSLLM=Llama3.1-8b2026.05 | 0.6471 | 0.6694 | 56.02 | 0.3173 | |
| NLI-MeanLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.6291 | 0.6899 | 64.98 | 0.4305 | |
| EntropyLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.6103 | 0.715 | 68.12 | 0.0952 | |
| EntropyLLM=Phi4-reasoning2026.05 | 0.6012 | 0.6702 | 57.09 | 0.052 | |
| Cos-MeanLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.5938 | 0.6919 | 64.19 | 0.2434 | |
| NLI-MaxLLM=Phi4-reasoning2026.05 | 0.5801 | 0.6637 | 69.61 | 0.2323 | |
| EntropyLLM=Llama3.1-8b2026.05 | 0.551 | 0.5413 | 55.86 | 0.2148 | |
| NLI-MeanLLM=Phi4-reasoning2026.05 | 0.544 | 0.6428 | 67.76 | 0.461 | |
| P(true)LLM=Llama3.1-8b2026.05 | 0.5228 | 0.5486 | 54.5 | 0.5357 | |
| P(true)LLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.5177 | 0.6492 | 66.06 | 0.6443 | |
| NLI-MeanLLM=Llama3.1-8b2026.05 | 0.5124 | 0.5535 | 53.14 | 0.3902 | |
| P(true)LLM=Phi4-reasoning2026.05 | 0.5086 | 0.6159 | 63.62 | 0.6203 | |
| Cos-MeanLLM=Llama3.1-8b2026.05 | 0.5044 | 0.5487 | 52.75 | 0.316 | |
| NLI-MaxLLM=Llama3.1-8b2026.05 | 0.4937 | 0.5485 | 53.47 | 0.2159 | |
| Cos-MeanLLM=Phi4-reasoning2026.05 | 0.4779 | 0.6492 | 63.88 | 0.2023 | |
| SL(norm)LLM=Llama3.1-8b2026.05 | 0.4005 | 0.4506 | 51.87 | 0.3168 | |
| Cos-MaxLLM=Llama3.1-8b2026.05 | 0.3836 | 0.4446 | 51.79 | 0.3687 | |
| Cos-MaxLLM=Phi4-reasoning2026.05 | 0.3454 | 0.5284 | 64.13 | 0.2614 | |
| Cos-MaxLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.339 | 0.5129 | 60.31 | 0.3066 | |
| LECOLLM=Llama3.1-8b2026.05 | 0.328 | 0.4365 | 49.21 | 0.3501 | |
| SL(norm)LLM=Phi4-reasoning2026.05 | 0.3198 | 0.4966 | 58.91 | 0.2751 | |
| LECOLLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.3116 | 0.5777 | 56.42 | 0.37 | |
| LECOLLM=Phi4-reasoning2026.05 | 0.276 | 0.4944 | 57.82 | 0.4254 | |
| SL(norm)LLM=DeepSeek-R1-Distill-Qwen-32B2026.05 | 0.2463 | 0.4894 | 60.49 | 0.2829 |