Multiple-choice question answering on MedMCQA (test)
73.7AccuracyGPT-4-base
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4-baseFew-shot evaluation protocol=5-shot2023.05 | 73.7 | |
| GPT-4Few-shot evaluation protocol=5-shot2023.05 | 72.4 | |
| Med-PaLM 2Prompting strategy=Ensemble Refinement (ER)2023.05 | 72.3 | |
| Med-PaLM 2Prompting strategy=best2023.05 | 72.3 | |
| Codex 5-shot CoTDate=2022, Few-shot shots=5, Chain-of-Thought (CoT)=true, Ensemble samples (k)=1002022.07 | 62.7 | |
| Flan-PaLMPrompting strategy=best2023.05 | 57.6 |