Confidence Calibration on SimpleQA
0.0386Brier ScoreProbe (train on TriviaQA)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Probe (train on TriviaQA)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.0386 | — | — | — | — | — | — | |
| P(True)LLM=Olmo-3-7B-Instruct, Cost=12026.02 | 0.0419 | — | — | — | — | — | — | |
| Probe (train on TriviaQA)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.06 | — | — | — | — | — | — | |
| Probe (train on TriviaQA)LLM=Qwen3-8B, Cost=< 12026.02 | 0.0638 | — | — | — | — | — | — | |
| GPT-5Zero-shot=true2025.12 | 0.178 | 0.498 | 0.838 | 72.9 | 0.098 | 0.53 | 43.8 | |
| Verbalized confidenceLLM=gpt-oss-20b, Cost=L2026.02 | 0.1957 | — | — | — | — | — | — | |
| Claude-sonnet-4.5Thinking=true, Zero-shot=true2025.12 | 0.214 | 0.456 | 0.783 | 65.8 | 0.2 | 0.614 | 30.8 | |
| Verbalized confidenceLLM=Olmo-3-7B-Instruct, Cost=L2026.02 | 0.2676 | — | — | — | — | — | — | |
| Claude-sonnet-4.5Thinking=false, Zero-shot=true2025.12 | 0.284 | 0.282 | 0.748 | 54.4 | 0.318 | 0.786 | 29.9 | |
| Uniform random baselineLLM=gpt-oss-20b, Cost=N/A2026.02 | 0.301 | — | — | — | — | — | — | |
| GLM-4.6Zero-shot=true2025.12 | 0.31 | 0.358 | 0.739 | 51.8 | 0.377 | 0.878 | 20.6 | |
| Uniform random baselineLLM=Qwen3-8B, Cost=N/A2026.02 | 0.3109 | — | — | — | — | — | — | |
| Uniform random baselineLLM=Olmo-3-7B-Instruct, Cost=N/A2026.02 | 0.3133 | — | — | — | — | — | — | |
| Grok-4Zero-shot=true2025.12 | 0.327 | 0.147 | 0.664 | 55.1 | 0.307 | 0.988 | 49.2 | |
| Confidence-BrierBase Model=Qwen3-4B-Instruct, Training Mode=confidence-Brier, Zero-shot=true2025.12 | 0.341 | 0.411 | 0.704 | 54 | 0.459 | 1.1 | 5.8 | |
| Gemini-2.5-proZero-shot=true2025.12 | 0.436 | 0.017 | 0.556 | 54.7 | 0.451 | 1.9 | 54.5 | |
| Probe (train on GSM8K)LLM=Qwen3-8B, Cost=< 12026.02 | 0.4451 | — | — | — | — | — | — | |
| Verbalized confidenceLLM=Qwen3-8B, Cost=L2026.02 | 0.4736 | — | — | — | — | — | — | |
| Probe (train on MATH)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.4846 | — | — | — | — | — | — | |
| Critic ValueBase Model=Qwen3-4B-Instruct, Training Mode=ppo-value, Zero-shot=true2025.12 | 0.525 | 0.211 | 0.734 | 23.6 | 0.665 | 1.51 | 3.1 | |
| Probe (train on GSM8K)LLM=Olmo-3-7B-Instruct, Cost=< 12026.02 | 0.5465 | — | — | — | — | — | — | |
| Probe (train on MATH)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.5871 | — | — | — | — | — | — | |
| P(True)LLM=Qwen3-8B, Cost=12026.02 | 0.6072 | — | — | — | — | — | — | |
| P(True)LLM=gpt-oss-20b, Cost=L2026.02 | 0.6151 | — | — | — | — | — | — | |
| Probe (train on GSM8K)LLM=gpt-oss-20b, Cost=< 12026.02 | 0.7048 | — | — | — | — | — | — | |
| Qwen3-4B-InstructZero-shot=true2025.12 | 0.762 | 0.066 | 0.561 | 14.2 | 0.821 | 2.55 | 6.2 | |
| Probe (train on MATH)LLM=Qwen3-8B, Cost=< 12026.02 | 0.8297 | — | — | — | — | — | — |