Calibration on TriviaQA
0.0845Brier ScoreProbe (train on TriviaQA)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Probe (train on TriviaQA)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.0845 | — | — | |
| Probe (train on TriviaQA)LLM=Qwen3-8B, Cost=< 12026.02 | 0.1079 | — | — | |
| Probe (train on TriviaQA)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.1113 | — | — | |
| Verbalized confidenceLLM=gpt-oss-20b, Cost=L2026.02 | 0.1266 | — | — | |
| Probe (train on MATH)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.1577 | — | — | |
| Temp. ScalingModel=Llama3.1-8B-Instruct, Unsupervised=false2026.01 | 0.1702 | 0.0226 | — | |
| BaseCal-ReEvalModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.1724 | 0.0309 | — | |
| Probe (train on GSM8K)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.1756 | — | — | |
| BaseCal-ReEvalModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.1789 | 0.112 | — | |
| BaseCal-ProjModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.185 | 0.0387 | — | |
| Temp. ScalingModel=Qwen2.5-7B-Instruct, Unsupervised=false2026.01 | 0.185 | 0.0895 | — | |
| Probe (train on GSM8K)LLM=Qwen3-8B, Cost=< 12026.02 | 0.1885 | — | — | |
| BaseCal-ProjModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.1923 | 0.1393 | — | |
| P(True)LLM=Olmo-3-7B-Instruct, Cost=12026.02 | 0.1933 | — | — | |
| Temp. ScalingModel=Olmo2-7B-Instruct, Unsupervised=false2026.01 | 0.1939 | 0.0286 | — | |
| BaseCal-ReEvalModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.1947 | 0.0269 | — | |
| BaseCal-ProjModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.1966 | 0.0314 | — | |
| VerbalizationModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.2046 | 0.1769 | — | |
| VanillaModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.209 | 0.1725 | — | |
| P(True)LLM=gpt-oss-20b, Cost=L2026.02 | 0.2101 | — | — | |
| P(true)Model=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.2134 | 0.2113 | — | |
| Semantic EntropyModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.2197 | 0.191 | — | |
| Semantic EntropyModel=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.2302 | 0.2443 | — | |
| P(true)Model=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.2393 | 0.2055 | — | |
| Verbalized confidenceLLM=Qwen3-8B, Cost=L2026.02 | 0.2431 | — | — | |
| VanillaModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.2465 | 0.2121 | — | |
| P(true)Model=Llama3.1-8B-Instruct, Unsupervised=true2026.01 | 0.2506 | 0.2476 | — | |
| Probe (train on MATH)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.255 | — | — | |
| Verbalized confidenceLLM=Olmo-3-7B-Instruct, Cost=L2026.02 | 0.2624 | — | — | |
| Uniform random baselineLLM=gpt-oss-20b, Cost=N/A2026.02 | 0.2639 | — | — | |
| Probe (train on GSM8K)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.2648 | — | — | |
| VerbalizationModel=Olmo2-7B-Instruct, Unsupervised=true2026.01 | 0.2721 | 0.2054 | — | |
| Uniform random baselineLLM=Olmo-3-7B-Instruct, Cost=N/A2026.02 | 0.2745 | — | — | |
| Uniform random baselineLLM=Qwen3-8B, Cost=N/A2026.02 | 0.2865 | — | — | |
| P(True)LLM=Qwen3-8B, Cost=12026.02 | 0.297 | — | — | |
| Probe (train on MATH)LLM=Qwen3-8B, Cost=< 12026.02 | 0.2977 | — | — | |
| VerbalizationModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.3031 | 0.2889 | — | |
| VanillaModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.3118 | 0.3406 | — | |
| Semantic EntropyModel=Qwen2.5-7B-Instruct, Unsupervised=true2026.01 | 0.3191 | 0.3583 | — | |
| CAGE-CALType=Trained2026.05 | — | 4.25 | 0.8612 | |
| Collab. Cal.Type=LLM-elicited2026.05 | — | 8.84 | 0.863 | |
| DiscoUQ-LLMType=Trained2026.05 | — | 6.88 | 0.8219 | |
| GraphCalType=Trained2026.05 | — | 11.05 | 0.8006 | |
| LLM-Cal (+topo)Type=LLM-elicited, Topology=Included2026.05 | — | 8.71 | 0.8573 | |
| LLM-Cal (no topo)Type=LLM-elicited, Topology=None2026.05 | — | 7.99 | 0.8574 | |
| Plurality shareType=Post-hoc plurality2026.05 | — | 11.43 | 0.8189 | |
| Plurality share + IsotonicType=Post-hoc plurality, Technique=Isotonic regression2026.05 | — | 8.11 | 0.819 | |
| Plurality share + PlattType=Post-hoc plurality, Technique=Platt scaling2026.05 | — | 7.96 | 0.8189 | |
| Plurality share + Scaling-bin.Type=Post-hoc plurality, Technique=Scaling-binning2026.05 | — | 9.58 | 0.813 | |
| Scalar + GBTType=Trained2026.05 | — | 9.13 | 0.7924 |