Confidence Estimation on MediTOD
68.7AUROCMedConf
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MedConfType=Self-Verbalized, Model=Llama-3.12026.01 | 68.7 | 0.943 | |
| MedConfModel=GPT-4.1, Type=Self-Verbalized2026.01 | 67.2 | 0.981 | |
| CEType=Self-Verbalized, Model=Llama-3.12026.01 | 63.6 | 0.895 | |
| SemSim (BERT)Type=Consistency-Level, Model=Llama-3.12026.01 | 62.8 | 0.491 | |
| CEModel=GPT-4.1, Type=Self-Verbalized2026.01 | 62.8 | 0.981 | |
| SemSim (BERT)Model=GPT-4.1, Type=Consistency-Level2026.01 | 59.3 | 0.829 | |
| Conditional PMIType=Token-Level, Model=Llama-3.12026.01 | 56 | 0.734 |