Medical Question Answering on MedQA US (4-option)
90.2AccuracyGPT-4 (Medprompt)
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4 (Medprompt)Prompting Strategy=Medprompt, Few-shot Count (k)=5, Ensemble steps=5x, Fine-tuning=false2023.11 | 90.2 | |
| GPT-4oPrompting (shots)=0-shot2024.10 | 89 | |
| GPT-4oPrompting (shots)=5-shot2024.10 | 89 | |
| Med-PaLM 2Strategy Selection=choose best, Fine-tuning=true2023.11 | 86.5 | |
| GPT-4Prompting Strategy=5-shot, Fine-tuning=false2023.11 | 81.4 | |
| GPT-4T (May 2024)Prompting (shots)=5-shot2024.10 | 81 | |
| GPT-4T (May 2024)Prompting (shots)=0-shot2024.10 | 78 | |
| Flan-PaLM 540BStrategy Selection=choose best, Fine-tuning=true2023.11 | 67.6 |