Capability Calibration on AIME 25
0.046Brier ScoreVerbalized confidence
Evaluation Results
| Method | Links | |
|---|---|---|
| Verbalized confidenceLLM=gpt-oss-20b, Cost=L2026.02 | 0.046 | |
| Probe (train on GSM8K)LLM=Qwen3-8B, Cost=< 12026.02 | 0.074 | |
| Probe (train on MATH)LLM=Qwen3-8B, Cost=< 12026.02 | 0.0831 | |
| P(True)LLM=gpt-oss-20b, Cost=L2026.02 | 0.1092 | |
| Probe (train on GSM8K)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.1213 | |
| Probe (train on TriviaQA)LLM=Qwen3-8B, Cost=< 12026.02 | 0.1286 | |
| Probe (train on MATH)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.1411 | |
| Probe (train on TriviaQA)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.1457 | |
| Probe (train on TriviaQA)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.1496 | |
| Probe (train on MATH)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.1644 | |
| P(True)LLM=Olmo-3-7B-Instruct, Cost=12026.02 | 0.1854 | |
| Verbalized confidenceLLM=Olmo-3-7B-Instruct, Cost=L2026.02 | 0.2002 | |
| Uniform random baselineLLM=gpt-oss-20b, Cost=N/A2026.02 | 0.2369 | |
| Uniform random baselineLLM=Olmo-3-7B-Instruct, Cost=N/A2026.02 | 0.2462 | |
| Probe (train on GSM8K)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.2482 | |
| Uniform random baselineLLM=Qwen3-8B, Cost=N/A2026.02 | 0.28 | |
| Verbalized confidenceLLM=Qwen3-8B, Cost=L2026.02 | 0.4443 | |
| P(True)LLM=Qwen3-8B, Cost=12026.02 | 0.4957 |