Calibration on MMLU
0.0559Brier ScoreVerbalized confidence
Evaluation Results
| Method | Links | ||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Verbalized confidenceLLM=gpt-oss-20b, Cost=L2026.02 | 0.0559 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on GSM8K)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.0686 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| P(True)LLM=gpt-oss-20b, Cost=L2026.02 | 0.0817 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on MATH)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.0922 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on TriviaQA)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.0977 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on MATH)LLM=Qwen3-8B, Cost=< 12026.02 | 0.1163 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on GSM8K)LLM=Qwen3-8B, Cost=< 12026.02 | 0.1176 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on GSM8K)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.12 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on MATH)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.1255 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Verbalized confidenceLLM=Qwen3-8B, Cost=L2026.02 | 0.1293 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on TriviaQA)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.13 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseCal-ReEvalModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.1465 | 0.0393 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseCal-ReEvalModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.1473 | 0.0375 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseCal-ProjModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.15 | 0.0336 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Verbalized confidenceLLM=Olmo-3-7B-Instruct, Cost=L2026.02 | 0.1561 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| P(True)LLM=Qwen3-8B, Cost=12026.02 | 0.1597 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DACAModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.1618 | 0.0703 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseCal-ProjModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.1662 | 0.0889 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseCal-ReEvalModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.1717 | 0.047 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseCal-ProjModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.1768 | 0.0525 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DACAModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.1804 | 0.0473 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DACAModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.1898 | 0.0555 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Probe (train on TriviaQA)LLM=Qwen3-8B, Cost=< 12026.02 | 0.2006 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Temp. ScalingModel=Llama3.1-8B-Instruct, Unsupervised=false2026.01 | 0.2021 | 0.0307 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.2152 | 0.1071 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| P(True)LLM=Olmo-3-7B-Instruct, Cost=12026.02 | 0.2164 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Temp. ScalingModel=Olmo2-7B-Instruct, Unsupervised=false2026.01 | 0.2275 | 0.1707 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Temp. ScalingModel=Qwen2.5-7B-Instruct, Unsupervised=false2026.01 | 0.2354 | 0.2261 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Llama-3.2-3B-Instruct, alpha=0.62026.04 | 0.242 | 0.087 | 48.9 | 8.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Llama-3.2-3B-Instruct, alpha=0.72026.04 | 0.242 | 0.079 | 53.7 | 8.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Llama-3.2-3B-Instruct, alpha=0.82026.04 | 0.242 | 0.074 | 57 | 8.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Llama-3.2-3B-Instruct, alpha=0.52026.04 | 0.243 | 0.095 | 44.5 | 8.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Llama-3.2-3B-Instruct, alpha=0.42026.04 | 0.245 | 0.104 | 39.4 | 7.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Llama-3.2-3B-Instruct, alpha=0.32026.04 | 0.246 | 0.112 | 34.7 | 7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VerbalizationModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.2546 | 0.1972 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| P(true)Model=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.2549 | 0.194 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Uniform random baselineLLM=Olmo-3-7B-Instruct, Cost=N/A2026.02 | 0.2565 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.2607 | 0.2569 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Qwen2.5-3B-Instruct, alpha=0.62026.04 | 0.263 | 0.188 | 33.1 | 15.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VerbalizationModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.2633 | 0.2011 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Qwen2.5-3B-Instruct, alpha=0.72026.04 | 0.265 | 0.197 | 29.9 | 14.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| baselineModel=Llama-3.2-3B-Instruct, alpha=–2026.04 | 0.265 | 0.171 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| mean ablationModel=Llama-3.2-3B-Instruct, alpha=–2026.04 | 0.269 | 0.166 | 3.1 | 1.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Qwen2.5-3B-Instruct, alpha=0.52026.04 | 0.271 | 0.196 | 30.3 | 12.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Qwen2.5-3B-Instruct, alpha=0.82026.04 | 0.274 | 0.217 | 22.7 | 11.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VerbalizationModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.2747 | 0.1533 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.2762 | 0.2465 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Qwen2.5-3B-Instruct, alpha=0.42026.04 | 0.284 | 0.23 | 18.2 | 8.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Semantic EntropyModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.2856 | 0.2858 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Uniform random baselineLLM=Qwen3-8B, Cost=N/A2026.02 | 0.2868 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| steeringModel=Qwen2.5-3B-Instruct, alpha=0.32026.04 | 0.294 | 0.25 | 11.1 | 5.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| mean ablationModel=Qwen2.5-3B-Instruct, alpha=–2026.04 | 0.298 | 0.261 | 7.3 | 3.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Uniform random baselineLLM=gpt-oss-20b, Cost=N/A2026.02 | 0.3018 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| P(true)Model=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.3079 | 0.2971 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| baselineModel=Qwen2.5-3B-Instruct, alpha=–2026.04 | 0.31 | 0.281 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| P(true)Model=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.3256 | 0.3204 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Semantic EntropyModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.3607 | 0.3132 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Semantic EntropyModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.3991 | 0.4085 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AverageIT=✗, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.0963 | 0.3927 | 0.3404 | 0.1625 | 0.3422 | 0.2877 | — | |
| AverageIT=✓, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.272 | 0.5485 | 0.453 | 0.2562 | 0.4998 | 0.378 | — | |
| AverageIT=✓, Chat=✓, Answer=Assistant2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.2877 | 0.6078 | 0.5014 | 0.2966 | 0.5468 | 0.4329 | — | |
| AverageIT=✓, Chat=✓, Answer=User2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.1897 | 0.4291 | 0.2404 | 0.2084 | 0.3521 | 0.1813 | — | |
| BCalibration Method=B2026.04 | — | — | 0.864 | 84.5 | 0 | 0.814 | 0 | 0 | 0.27 | 0 | 0.698 | 0 | — | — | — | — | — | — | — | — | |
| BaseModel=LLAMA3-8B, Temperature Scaling=false2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.281 | |
| BaseModel=LLAMA3-8B, Temperature Scaling=true2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.1 | |
| BaseModel=QWEN2-7B, Temperature Scaling=false2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.273 | |
| BaseModel=QWEN2-7B, Temperature Scaling=true2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.133 | |
| Gemma 3 (27B)IT=✗, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.0428 | 0.5213 | 0.3356 | 0.1798 | 0.4465 | 0.2942 | — | |
| Gemma 3 (27B)IT=✓, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.2395 | 0.5465 | 0.5256 | 0.2366 | 0.4785 | 0.452 | — | |
| Gemma 3 (27B)IT=✓, Chat=✓, Answer=Assistant2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.3179 | 0.6771 | 0.456 | 0.3171 | 0.6461 | 0.3813 | — | |
| Gemma 3 (27B)IT=✓, Chat=✓, Answer=User2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.1919 | 0.3621 | 0.265 | 0.1911 | 0.3011 | 0.2063 | — | |
| Gemma 3 (4B)IT=✗, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.0209 | 0.4558 | 0.3452 | 0.1882 | 0.393 | 0.3062 | — | |
| Gemma 3 (4B)IT=✓, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.5779 | 0.7132 | 0.5418 | 0.5645 | 0.6929 | 0.4731 | — | |
| Gemma 3 (4B)IT=✓, Chat=✓, Answer=Assistant2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.6546 | 0.713 | 0.5477 | 0.6517 | 0.6953 | 0.4959 | — | |
| Gemma 3 (4B)IT=✓, Chat=✓, Answer=User2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.4026 | 0.592 | 0.2797 | 0.3987 | 0.5463 | 0.2499 | — | |
| GIRBCalibration Method=GIRB2026.04 | — | — | 1.145 | 90.1 | 0 | 1.167 | 3 | 4 | 0.531 | 3 | 0.936 | 2.5 | — | — | — | — | — | — | — | — | |
| HS-QABCalibration Method=HS-QAB2026.04 | — | — | 0.739 | 64.8 | 0 | 0.968 | 0 | 0 | 0.039 | 0 | 0.599 | 0 | — | — | — | — | — | — | — | — | |
| Llama 3.1 (70B)IT=✗, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.0267 | 0.1733 | 0.2745 | 0.0992 | 0.1405 | 0.1983 | — | |
| Llama 3.1 (70B)IT=✓, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.0744 | 0.2135 | 0.3151 | 0.103 | 0.1739 | 0.2288 | — | |
| Llama 3.1 (70B)IT=✓, Chat=✓, Answer=Assistant2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.113 | 0.4137 | 0.4511 | 0.1825 | 0.25 | 0.3818 | — | |
| Llama 3.1 (70B)IT=✓, Chat=✓, Answer=User2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.022 | 0.3937 | 0.4261 | 0.1155 | 0.228 | 0.3528 | — | |
| Llama 3.1 (8B)IT=✗, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.191 | 0.3483 | 0.3161 | 0.193 | 0.3174 | 0.2872 | — | |
| Llama 3.1 (8B)IT=✓, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.2796 | 0.649 | 0.3767 | 0.2315 | 0.5992 | 0.3252 | — | |
| Llama 3.1 (8B)IT=✓, Chat=✓, Answer=Assistant2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.2329 | 0.5133 | 0.4414 | 0.2236 | 0.45 | 0.3707 | — | |
| Llama 3.1 (8B)IT=✓, Chat=✓, Answer=User2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.1609 | 0.4373 | 0.2064 | 0.1916 | 0.379 | 0.1547 | — | |
| NoneCalibration Method=None2026.04 | — | — | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | — | — | — | — | — | — | — | — | |
| PPT (Distribution)Model=LLAMA3-8B, Temperature Scaling=false2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.112 | |
| PPT (Distribution)Model=LLAMA3-8B, Temperature Scaling=true2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.069 | |
| PPT (Distribution)Model=QWEN2-7B, Temperature Scaling=false2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.18 | |
| PPT (Distribution)Model=QWEN2-7B, Temperature Scaling=true2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.095 | |
| PPT (Mean)Model=LLAMA3-8B, Temperature Scaling=false2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.216 | |
| PPT (Mean)Model=QWEN2-7B, Temperature Scaling=false2026.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.254 | |
| QABCalibration Method=QAB2026.04 | — | — | 1.061 | 87.3 | 0 | 1.13 | 0 | 0 | 0.429 | 1 | 0.873 | 0.2 | — | — | — | — | — | — | — | — | |
| Qwen3 (30B)IT=✗, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.1326 | 0.4037 | 0.3589 | 0.1415 | 0.3377 | 0.2806 | — | |
| Qwen3 (30B)IT=✓, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.234 | 0.5835 | 0.4565 | 0.1829 | 0.5001 | 0.365 | — | |
| Qwen3 (30B)IT=✓, Chat=✓, Answer=Assistant2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.1935 | 0.6759 | 0.5674 | 0.1902 | 0.6368 | 0.4916 | — | |
| Qwen3 (30B)IT=✓, Chat=✓, Answer=User2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.1485 | 0.3289 | 0.1314 | 0.1502 | 0.2828 | 0.0636 | — | |
| Qwen3 (4B)IT=✗, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.1635 | 0.4539 | 0.4123 | 0.1732 | 0.4182 | 0.3598 | — | |
| Qwen3 (4B)IT=✓, Chat=✗, Answer=N/A2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.2268 | 0.585 | 0.5021 | 0.2186 | 0.5544 | 0.4241 | — | |
| Qwen3 (4B)IT=✓, Chat=✓, Answer=Assistant2026.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | 0.214 | 0.6535 | 0.5445 | 0.2145 | 0.6026 | 0.4762 | — |