Uncertainty Estimation on TriviaQA (test)
87.91AUROCACT-ViT
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ACT-ViTTrain Dataset=HotpotQA, Model=Mistral-7B-Instruct-v0.32026.03 | 87.91 | — | — | — | |
| CAGE-CALRollouts=3, Topologies=52026.05 | 86.12 | — | — | 93.02 | |
| TSSC Measure=IC-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 85.7 | — | — | — | |
| TSSC Measure=E-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 84.9 | — | — | — | |
| TSSC Measure=G-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 84.8 | — | — | — | |
| TSSC Measure=L-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 84.4 | — | — | — | |
| ATSSC Measure=E-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.4 | — | — | — | |
| ATSSC Measure=G-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.4 | — | — | — | |
| SESC Measure=E-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.3 | — | — | — | |
| SESC Measure=L-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.3 | — | — | — | |
| SESC Measure=IC-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.3 | — | — | — | |
| SESC Measure=G-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.3 | — | — | — | |
| PlattSC Measure=E-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.2 | — | — | — | |
| PlattSC Measure=L-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.2 | — | — | — | |
| PlattSC Measure=IC-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.2 | — | — | — | |
| PlattSC Measure=G-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.2 | — | — | — | |
| ATSSC Measure=L-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.2 | — | — | — | |
| ATSSC Measure=IC-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 83.2 | — | — | — | |
| ACT-ViTTrain Dataset=TriviaQA, Model=Mistral-7B-Instruct-v0.32026.03 | 82.66 | — | — | — | |
| SETarget Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 82.12 | 12.76 | 1.2 | — | |
| SARTarget Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 81.9 | 13.76 | 0.98 | — | |
| Plurality voteRollouts=3, Topologies=52026.05 | 81.89 | — | — | 93.09 | |
| SARModel=LLaMA-7b, Rouge-L threshold=0.32023.07 | 81.8 | — | — | — | |
| SEModel=LLaMA-7b, Rouge-L threshold=0.32023.07 | 81.4 | — | — | — | |
| SETarget Model=OPT-6.7B, Setting=+Corrector2025.05 | 81 | 27.05 | 0.34 | — | |
| SETarget Model=LLaMA-3-8B-Instruct, Setting=Vanilla2025.05 | 80.92 | 13.07 | — | — | |
| SARTarget Model=LLaMA-3-8B-Instruct, Setting=Vanilla2025.05 | 80.92 | 16.17 | — | — | |
| SETarget Model=OPT-6.7B, Setting=Vanilla2025.05 | 80.66 | 36.64 | — | — | |
| SAR-tTarget Model=OPT-6.7B, Setting=+Corrector2025.05 | 80.21 | 37.85 | 1.97 | — | |
| Answer entropyRollouts=3, Topologies=52026.05 | 80.19 | — | — | 92.82 | |
| DAERollouts=3, Topologies=52026.05 | 80.19 | — | — | 92.82 | |
| SENTSARModel=LLaMA-7b, Rouge-L threshold=0.32023.07 | 80 | — | — | — | |
| LN-PETarget Model=OPT-6.7B, Setting=+Corrector2025.05 | 79.93 | 20.8 | 1.57 | — | |
| SAR-tTarget Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 79.93 | 13.7 | 0.38 | — | |
| SAR-tTarget Model=LLaMA-3-8B-Instruct, Setting=Vanilla2025.05 | 79.55 | 16.4 | — | — | |
| TOKENSARModel=LLaMA-7b, Rouge-L threshold=0.32023.07 | 79.3 | — | — | — | |
| SARTarget Model=OPT-6.7B, Setting=+Corrector2025.05 | 78.67 | 31.02 | 3.35 | — | |
| LN-PETarget Model=OPT-6.7B, Setting=Vanilla2025.05 | 78.37 | 32.29 | — | — | |
| LN-PEModel=LLaMA-7b, Rouge-L threshold=0.32023.07 | 78.3 | — | — | — | |
| SAR-tTarget Model=OPT-6.7B, Setting=Vanilla2025.05 | 78.24 | 40.14 | — | — | |
| SAR-sTarget Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 77.09 | 20 | 2.95 | — | |
| SEModel=LLaMA-13b, Rouge-L threshold=0.32023.07 | 75.8 | — | — | — | |
| BaseSC Measure=E-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 75.5 | — | — | — | |
| BaseSC Measure=L-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 75.5 | — | — | — | |
| BaseSC Measure=G-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 75.5 | — | — | — | |
| BaseSC Measure=IC-SC, Model=Qwen, Evaluation Protocol=SE_vanilla (greedy-decoding)2026.04 | 75.4 | — | — | — | |
| SARTarget Model=OPT-6.7B, Setting=Vanilla2025.05 | 75.32 | 40.61 | — | — | |
| SENTSARModel=LLaMA-13b, Rouge-L threshold=0.32023.07 | 74.9 | — | — | — | |
| SARLLM Model=Vicuna-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 74.9 | — | — | — | |
| VCTarget Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 74.89 | 16.78 | 12.55 | — | |
| LN-PETarget Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 74.79 | 11.53 | 2.24 | — | |
| SENTSARLLM Model=Vicuna-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 74.5 | — | — | — | |
| SARLLM Model=WizardLM-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 74.4 | — | — | — | |
| SENTSARLLM Model=WizardLM-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 74.3 | — | — | — | |
| LN-PETarget Model=LLaMA-3-8B-Instruct, Setting=Vanilla2025.05 | 72.55 | 14.31 | — | — | |
| P(True)Target Model=OPT-6.7B, Setting=+Corrector2025.05 | 72.29 | 32.63 | 5.84 | — | |
| P(True)Target Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 72.29 | 19.84 | 15.15 | — | |
| SARModel=LLaMA-13b, Rouge-L threshold=0.32023.07 | 71.6 | — | — | — | |
| SENTSARLLM Model=Vicuna-33b, Number of generations=5, Rouge-L threshold=0.52023.07 | 71.5 | — | — | — | |
| PEModel=LLaMA-7b, Rouge-L threshold=0.32023.07 | 71.3 | — | — | — | |
| MATURollouts=3, Topologies=52026.05 | 71.23 | — | — | 89.73 | |
| SignaturesTrain Dataset=IMDB, Model=Mistral-7B-Instruct-v0.32026.03 | 71.19 | — | — | — | |
| SARLLM Model=Vicuna-33b, Number of generations=5, Rouge-L threshold=0.52023.07 | 71 | — | — | — | |
| VCTarget Model=OPT-6.7B, Setting=+Corrector2025.05 | 70.55 | 27.61 | 10.15 | — | |
| SARLLM Model=LLaMA-2-13b-chat, Number of generations=5, Rouge-L threshold=0.52023.07 | 70.4 | — | — | — | |
| SAR-sTarget Model=LLaMA-3-8B-Instruct, Setting=Vanilla2025.05 | 69.87 | 23.17 | — | — | |
| Avg-logprobRollouts=3, Topologies=52026.05 | 69.83 | — | — | 86.96 | |
| LSTarget Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 69.82 | 7.41 | 50.25 | — | |
| SENTSARLLM Model=LLaMA-2-13b-chat, Number of generations=5, Rouge-L threshold=0.52023.07 | 69.8 | — | — | — | |
| PETarget Model=LLaMA-3-8B-Instruct, Setting=+Corrector2025.05 | 69.76 | 17.24 | 5.25 | — | |
| TOKENSARLLM Model=Vicuna-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 69.2 | — | — | — | |
| PELLM Model=Vicuna-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 69 | — | — | — | |
| P(True)Target Model=OPT-6.7B, Setting=Vanilla2025.05 | 66.74 | 45 | — | — | |
| PETarget Model=OPT-6.7B, Setting=+Corrector2025.05 | 66.62 | 20.28 | 10.25 | — | |
| TOKENSARLLM Model=WizardLM-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 65.7 | — | — | — | |
| LSModel=LLaMA-13b, Rouge-L threshold=0.32023.07 | 65.5 | — | — | — | |
| TOKENSARLLM Model=LLaMA-2-13b-chat, Number of generations=5, Rouge-L threshold=0.52023.07 | 65.4 | — | — | — | |
| TOKENSARLLM Model=Vicuna-33b, Number of generations=5, Rouge-L threshold=0.52023.07 | 65.2 | — | — | — | |
| LSTarget Model=OPT-6.7B, Setting=+Corrector2025.05 | 65.11 | 41.76 | 18.62 | — | |
| SELLM Model=Vicuna-33b, Number of generations=5, Rouge-L threshold=0.52023.07 | 65.1 | — | — | — | |
| PELLM Model=WizardLM-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 64.7 | — | — | — | |
| PELLM Model=LLaMA-2-13b-chat, Number of generations=5, Rouge-L threshold=0.52023.07 | 64.7 | — | — | — | |
| PETarget Model=LLaMA-3-8B-Instruct, Setting=Vanilla2025.05 | 64.52 | 21.38 | — | — | |
| PELLM Model=Vicuna-33b, Number of generations=5, Rouge-L threshold=0.52023.07 | 64.4 | — | — | — | |
| LN-PELLM Model=Vicuna-33b, Number of generations=5, Rouge-L threshold=0.52023.07 | 63.9 | — | — | — | |
| TOKENSARModel=LLaMA-13b, Rouge-L threshold=0.32023.07 | 63.5 | — | — | — | |
| SELLM Model=WizardLM-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 63.4 | — | — | — | |
| SELLM Model=Vicuna-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 63 | — | — | — | |
| LN-PEModel=LLaMA-13b, Rouge-L threshold=0.32023.07 | 62.7 | — | — | — | |
| LN-PELLM Model=Vicuna-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 62.4 | — | — | — | |
| VCTarget Model=LLaMA-3-8B-Instruct, Setting=Vanilla2025.05 | 62.34 | 23.41 | — | — | |
| SELLM Model=LLaMA-2-13b-chat, Number of generations=5, Rouge-L threshold=0.52023.07 | 62.2 | — | — | — | |
| ACT-ViTTrain Dataset=IMDB, Model=Mistral-7B-Instruct-v0.32026.03 | 62.01 | — | — | — | |
| LN-PELLM Model=WizardLM-13b, Number of generations=5, Rouge-L threshold=0.52023.07 | 61.5 | — | — | — | |
| LN-PELLM Model=LLaMA-2-13b-chat, Number of generations=5, Rouge-L threshold=0.52023.07 | 61.5 | — | — | — | |
| VCTarget Model=OPT-6.7B, Setting=Vanilla2025.05 | 60.41 | 49.13 | — | — | |
| SignaturesTrain Dataset=HotpotQA, Model=Mistral-7B-Instruct-v0.32026.03 | 58.42 | — | — | — | |
| P(True)Target Model=LLaMA-3-8B-Instruct, Setting=Vanilla2025.05 | 57.14 | 24.67 | — | — | |
| LSLLM Model=Vicuna-33b, Number of generations=5, Rouge-L threshold=0.52023.07 | 56.5 | — | — | — | |
| PETarget Model=OPT-6.7B, Setting=Vanilla2025.05 | 56.36 | 42.39 | — | — |