Medical Question Answering on Standard Questions
20.25Baseline Wins (%)ToT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ToTEvaluator=GPT-42024.12 | 20.25 | 40.69 | 39.06 | |
| Self-AlignEvaluator=GPT-42024.12 | 16.35 | 42.18 | 41.47 | |
| CoT-SC (3)Evaluator=GPT-42024.12 | 15.67 | 47.36 | 36.97 | |
| ToTEvaluator=Human2024.12 | 14.29 | 34.29 | 51.43 | |
| Few-shot (2)Evaluator=GPT-42024.12 | 12.77 | 54.98 | 32.25 | |
| CoTEvaluator=GPT-42024.12 | 12.48 | 67.27 | 20.25 | |
| CoT-SC (3)Evaluator=Human2024.12 | 11.43 | 31.43 | 57.14 | |
| (50) casesEvaluator=Human2024.12 | 11.23 | 20.72 | 68.05 | |
| (50) casesEvaluator=GPT-42024.12 | 10.38 | 18.15 | 71.47 | |
| CoTEvaluator=Human2024.12 | 9.35 | 45.26 | 45.39 | |
| Few-shot (2)Evaluator=Human2024.12 | 6.94 | 29.41 | 63.65 | |
| Self-AlignEvaluator=Human2024.12 | 6.06 | 34.38 | 59.56 |