Medical Reasoning on MedCaseReasoning (test)
72.5AccuracyGPT-5-mini
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-5-miniBackbone Model=GPT-5-mini, Confidence Estimation Strategy=Base2025.11 | 72.5 | — | — | — | |
| GPT-5-mini + Verbalized Conf.Backbone Model=GPT-5-mini, Confidence Estimation Strategy=Verbalized Conf.2025.11 | 72 | 73.7 | 0.192 | 0.12 | |
| GPT-5-mini + Verbalized Top-kBackbone Model=GPT-5-mini, Confidence Estimation Strategy=Verbalized Top-k2025.11 | 70.2 | 70.2 | 0.195 | 0.084 | |
| GPT-5-mini + Verbalized Probability DistributionBackbone Model=GPT-5-mini, Confidence Estimation Strategy=Verbalized Probability Distribution2025.11 | 69.1 | 72.3 | 0.196 | 0.103 | |
| GPT-4.1 + Verbalized Conf.Backbone=GPT-4.1, Inference Strategy=+ Verbalized Conf.2025.11 | 66.5 | 62.9 | 0.284 | 0.263 | |
| GPT-4.1 + Verbalized Distrib.Backbone=GPT-4.1, Inference Strategy=+ Verbalized Probability Distribution2025.11 | 65.5 | 66.3 | 0.211 | 0.03 | |
| GPT-4.1Backbone=GPT-4.1, Inference Strategy=Base2025.11 | 64.4 | — | — | — | |
| GPT-4.1 + Verbalized Top-kBackbone=GPT-4.1, Inference Strategy=+ Verbalized Top-k2025.11 | 62.3 | 65.4 | 0.248 | 0.165 | |
| DeepSeek-V3 + Verbalized Conf.Backbone=DeepSeek-V3, Inference Strategy=+ Verbalized Conf.2025.11 | 57.5 | 67.3 | 0.308 | 0.283 | |
| DeepSeek-V3Backbone=DeepSeek-V3, Inference Strategy=Base2025.11 | 56.8 | — | — | — | |
| DeepSeek-V3 + Verbalized Distrib.Backbone=DeepSeek-V3, Inference Strategy=+ Verbalized Probability Distribution2025.11 | 55.1 | 68.4 | 0.227 | 0.071 | |
| DeepSeek-V3 + Verbalized Top-kBackbone=DeepSeek-V3, Inference Strategy=+ Verbalized Top-k2025.11 | 54.3 | 64.4 | 0.292 | 0.242 | |
| Qwen3-30B-A3B-Instruct + Verbalized Top-kBackbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ Verbalized Top-k2025.11 | 44.9 | 58.3 | 0.472 | 0.48 | |
| Qwen3-30B-A3B-Instruct + Verbalized Conf.Backbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ Verbalized Conf.2025.11 | 44.3 | 58.7 | 0.506 | 0.512 | |
| Qwen3-30B-A3B-Instruct + Verbalized Distrib.Backbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ Verbalized Probability Distribution2025.11 | 42.9 | 67.3 | 0.289 | 0.252 | |
| Qwen3-30B-A3B-InstructBackbone=Qwen3-30B-A3B-Instruct, Inference Strategy=Base2025.11 | 40.4 | — | — | — | |
| Qwen3-30B-A3B-Instruct + LogitBackbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ Logit2025.11 | 40.4 | 56.7 | 0.381 | 0.338 | |
| Qwen3-30B-A3B-Instruct + p(True)Backbone=Qwen3-30B-A3B-Instruct, Inference Strategy=+ p(True)2025.11 | 40.4 | 59.5 | 0.561 | 0.563 | |
| Qwen3-4B-ThinkingBackbone Model=Qwen3-4B-Thinking, Confidence Estimation Strategy=Base2025.11 | 31.7 | — | — | — | |
| Qwen3-4B-Instruct + Verbalized Conf.Backbone=Qwen3-4B-Instruct, Inference Strategy=+ Verbalized Conf.2025.11 | 31.5 | 64.7 | 0.585 | 0.612 | |
| Qwen3-4B-Thinking + Verbalized Conf.Backbone Model=Qwen3-4B-Thinking, Confidence Estimation Strategy=Verbalized Conf.2025.11 | 31.5 | 63.9 | 0.509 | 0.546 | |
| Qwen3-4B-Instruct + Verbalized Distrib.Backbone=Qwen3-4B-Instruct, Inference Strategy=+ Verbalized Probability Distribution2025.11 | 31.1 | 64.8 | 0.33 | 0.34 | |
| Qwen3-4B-Thinking + Verbalized Top-kBackbone Model=Qwen3-4B-Thinking, Confidence Estimation Strategy=Verbalized Top-k2025.11 | 30.9 | 65.9 | 0.426 | 0.471 | |
| Qwen3-4B-Instruct + Verbalized Top-kBackbone=Qwen3-4B-Instruct, Inference Strategy=+ Verbalized Top-k2025.11 | 30.1 | 62.3 | 0.527 | 0.568 | |
| Qwen3-4B-InstructBackbone=Qwen3-4B-Instruct, Inference Strategy=Base2025.11 | 29.8 | — | — | — | |
| Qwen3-4B-Instruct + LogitBackbone=Qwen3-4B-Instruct, Inference Strategy=+ Logit2025.11 | 29.8 | 49.2 | 0.519 | 0.513 | |
| Qwen3-4B-Instruct + p(True)Backbone=Qwen3-4B-Instruct, Inference Strategy=+ p(True)2025.11 | 29.8 | 65.1 | 0.661 | 0.665 | |
| Qwen3-4B-Thinking + Verbalized Probability DistributionBackbone Model=Qwen3-4B-Thinking, Confidence Estimation Strategy=Verbalized Probability Distribution2025.11 | 29 | 66.2 | 0.356 | 0.394 |