Question Answering on HotpotQA (test) (ACC, AUROC, Brier, ECE Metrics)
68.6AccuracyGPT-4.1
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-4.1Backbone=GPT-4.1, Inference Strategy=Base2025.11 | 68.6 | — | — | — | |
| GPT-4.1 + Verbalized Distrib.Backbone=GPT-4.1, Inference Strategy=+ Verbalized Probability Distribution2025.11 | 68.2 | 77.2 | 20.8 | 17 | |
| GPT-5-mini + Verbalized Conf.Backbone Model=GPT-5-mini, Confidence Estimation Strategy=Verbalized Conf.2025.11 | 68.2 | 84 | 15.5 | 8.4 | |
| GPT-4.1 + Verbalized Conf.Backbone=GPT-4.1, Inference Strategy=+ Verbalized Conf.2025.11 | 67.9 | 72.8 | 29.1 | 29 | |
| GPT-4.1 + Verbalized Top-kBackbone=GPT-4.1, Inference Strategy=+ Verbalized Top-k2025.11 | 67.3 | 75.5 | 23.2 | 20.6 | |
| GPT-5-miniBackbone Model=GPT-5-mini, Confidence Estimation Strategy=Base2025.11 | 66.9 | — | — | — | |
| GPT-5-mini + Verbalized Probability DistributionBackbone Model=GPT-5-mini, Confidence Estimation Strategy=Verbalized Probability Distribution2025.11 | 66.8 | 86.2 | 15.1 | 8.2 | |
| GPT-5-mini + Verbalized Top-kBackbone Model=GPT-5-mini, Confidence Estimation Strategy=Verbalized Top-k2025.11 | 65.9 | 87 | 17.4 | 13.7 | |
| DeepSeek-V3 + Verbalized Distrib.Backbone=DeepSeek-V3, Inference Strategy=+ Verbalized Probability Distribution2025.11 | 59.4 | 77.7 | 20.3 | 11.6 | |
| DeepSeek-V3 + Verbalized Conf.Backbone=DeepSeek-V3, Inference Strategy=+ Verbalized Conf.2025.11 | 59.1 | 75 | 29.3 | 28.7 | |
| DeepSeek-V3Backbone=DeepSeek-V3, Inference Strategy=Base2025.11 | 58.3 | — | — | — | |
| DeepSeek-V3 + Verbalized Top-kBackbone=DeepSeek-V3, Inference Strategy=+ Verbalized Top-k2025.11 | 57.2 | 69.8 | 27.6 | 24.1 | |
| Qwen3-30B-A3B-Instruct + Verbalized Top-kBackbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ Verbalized Top-k2025.11 | 43.9 | 68.4 | 46.3 | 47.4 | |
| Qwen3-30B-A3B-Instruct + Verbalized Distrib.Backbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ Verbalized Probability Distribution2025.11 | 43.8 | 68.9 | 31.8 | 29.5 | |
| Qwen3-30B-A3B-InstructBackbone=Qwen3-30B-A3B-Instruct, Inference Strategy=Base2025.11 | 43.7 | — | — | — | |
| Qwen3-30B-A3B-Instruct + LogitBackbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ Logit2025.11 | 43.7 | 59.6 | 48.5 | 48.6 | |
| Qwen3-30B-A3B-Instruct + p(True)Backbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ p(True)2025.11 | 43.7 | 70.2 | 46 | 46.2 | |
| Qwen3-30B-A3B-Instruct + Verbalized Conf.Backbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ Verbalized Conf.2025.11 | 42.6 | 66.4 | 51.8 | 52.6 | |
| Qwen3-4B-Instruct + Verbalized Distrib.Backbone=Qwen3-4B-Instruct, Inference Strategy=+ Verbalized Probability Distribution2025.11 | 33.6 | 71.4 | 33.6 | 32.2 | |
| Qwen3-4B-Thinking + Verbalized Top-kBackbone Model=Qwen3-4B-Thinking, Confidence Estimation Strategy=Verbalized Top-k2025.11 | 33 | 77.6 | 37.6 | 43.2 | |
| Qwen3-4B-Instruct + Verbalized Top-kBackbone=Qwen3-4B-Instruct, Inference Strategy=+ Verbalized Top-k2025.11 | 31.9 | 65.6 | 55 | 57.2 | |
| Qwen3-4B-Thinking + Verbalized Conf.Backbone Model=Qwen3-4B-Thinking, Confidence Estimation Strategy=Verbalized Conf.2025.11 | 31.7 | 76.5 | 33.3 | 38.5 | |
| Qwen3-4B-Thinking + Verbalized Probability DistributionBackbone Model=Qwen3-4B-Thinking, Confidence Estimation Strategy=Verbalized Probability Distribution2025.11 | 31.7 | 78.1 | 20.3 | 16.9 | |
| Qwen3-4B-InstructBackbone=Qwen3-4B-Instruct, Inference Strategy=Base2025.11 | 30.3 | — | — | — | |
| Qwen3-4B-Instruct + LogitBackbone=Qwen3-4B-Instruct, Inference Strategy=+ Logit2025.11 | 30.3 | 58.6 | 63.9 | 64.7 | |
| Qwen3-4B-Instruct + p(True)Backbone=Qwen3-4B-Instruct, Inference Strategy=+ p(True)2025.11 | 30.3 | 73.3 | 48.7 | 49 | |
| Qwen3-4B-ThinkingBackbone Model=Qwen3-4B-Thinking, Confidence Estimation Strategy=Base2025.11 | 29.9 | — | — | — | |
| Qwen3-4B-Instruct + Verbalized Conf.Backbone=Qwen3-4B-Instruct, Inference Strategy=+ Verbalized Conf.2025.11 | 29 | 64.2 | 63.2 | 64.7 |