Capability Calibration on GSM8K
0.0268Brier ScoreVerbalized confidence
Evaluation Results
| Method | Links | |
|---|---|---|
| Verbalized confidenceLLM=gpt-oss-20b, Cost=L2026.02 | 0.0268 | |
| Probe (train on GSM8K)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.0289 | |
| P(True)LLM=gpt-oss-20b, Cost=L2026.02 | 0.0306 | |
| Probe (train on MATH)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.0332 | |
| Probe (train on GSM8K)LLM=Qwen3-8B, Cost=< 12026.02 | 0.0368 | |
| Probe (train on GSM8K)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.037 | |
| Probe (train on MATH)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.0388 | |
| Probe (train on MATH)LLM=Qwen3-8B, Cost=< 12026.02 | 0.0408 | |
| Verbalized confidenceLLM=Qwen3-8B, Cost=L2026.02 | 0.0461 | |
| Verbalized confidenceLLM=Olmo-3-7B-Instruct, Cost=L2026.02 | 0.0462 | |
| P(True)LLM=Qwen3-8B, Cost=12026.02 | 0.0482 | |
| Probe (train on TriviaQA)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.078 | |
| Probe (train on TriviaQA)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.118 | |
| P(True)LLM=Olmo-3-7B-Instruct, Cost=12026.02 | 0.1282 | |
| Uniform random baselineLLM=Olmo-3-7B-Instruct, Cost=N/A2026.02 | 0.3119 | |
| Uniform random baselineLLM=Qwen3-8B, Cost=N/A2026.02 | 0.3144 | |
| Probe (train on TriviaQA)LLM=Qwen3-8B, Cost=< 12026.02 | 0.3177 | |
| Uniform random baselineLLM=gpt-oss-20b, Cost=N/A2026.02 | 0.3195 |