Question Answering on MedMCQA (dev)
0.791AccuracyGPT-4 (Medprompt)
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4 (Medprompt)Prompting Strategy=Medprompt, Few-shot Count (k)=5, Ensemble steps=5x, Fine-tuning=false2023.11 | 0.791 | |
| GPT-4oPrompting (shots)=5-shot2024.10 | 0.79 | |
| GPT-4oPrompting (shots)=0-shot2024.10 | 0.77 | |
| GPT-4Prompting Strategy=5-shot, Fine-tuning=false2023.11 | 0.724 | |
| Med-PaLM 2Strategy Selection=choose best, Fine-tuning=true2023.11 | 0.723 | |
| GPT-4T (May 2024)Prompting (shots)=5-shot2024.10 | 0.72 | |
| GPT-4T (May 2024)Prompting (shots)=0-shot2024.10 | 0.7 | |
| Flan-PaLM 540BStrategy Selection=choose best, Fine-tuning=true2023.11 | 0.576 | |
| GalacticaModel Size=120B, Shots=0, Domain=in-domain2022.11 | 0.529 | |
| BLOOMShots=5, Domain=in-domain2022.11 | 0.325 | |
| OPTShots=5, Domain=in-domain2022.11 | 0.296 |