Capability Calibration on GPQA
0.1174Brier ScoreVerbalized confidence
Evaluation Results
| Method | Links | |
|---|---|---|
| Verbalized confidenceLLM=gpt-oss-20b, Cost=L2026.02 | 0.1174 | |
| Probe (train on TriviaQA)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.1242 | |
| Probe (train on MATH)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.1295 | |
| Probe (train on MATH)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.1363 | |
| Probe (train on TriviaQA)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.1533 | |
| P(True)LLM=Olmo-3-7B-Instruct, Cost=12026.02 | 0.1553 | |
| Probe (train on TriviaQA)LLM=Qwen3-8B, Cost=< 12026.02 | 0.1556 | |
| Probe (train on GSM8K)LLM=Qwen3-8B, Cost=< 12026.02 | 0.1556 | |
| Probe (train on GSM8K)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.1628 | |
| Probe (train on MATH)LLM=Qwen3-8B, Cost=< 12026.02 | 0.1811 | |
| Probe (train on GSM8K)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.201 | |
| P(True)LLM=gpt-oss-20b, Cost=L2026.02 | 0.2082 | |
| Uniform random baselineLLM=Qwen3-8B, Cost=N/A2026.02 | 0.2113 | |
| Uniform random baselineLLM=Olmo-3-7B-Instruct, Cost=N/A2026.02 | 0.2125 | |
| Uniform random baselineLLM=gpt-oss-20b, Cost=N/A2026.02 | 0.2388 | |
| Verbalized confidenceLLM=Olmo-3-7B-Instruct, Cost=L2026.02 | 0.2742 | |
| Verbalized confidenceLLM=Qwen3-8B, Cost=L2026.02 | 0.2773 | |
| P(True)LLM=Qwen3-8B, Cost=12026.02 | 0.3448 |