Question Answering on MedQA, MedQA-5, PubMedQA, and MedMCQA
60.88MedQA AccuracyLlama + SFTMix
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Llama + SFTMixLLM=Llama, Training Dataset=MedAlpaca-263K, Training Recipe=SFTMix, Evaluation Protocol=three-shot setting2024.10 | 60.88 | 55.38 | 77.8 | 54.15 | 62.05 | |
| LlamaLLM=Llama, Training Dataset=MedAlpaca-263K, Evaluation Protocol=three-shot setting2024.10 | 59.68 | 53.23 | 73.4 | 52.79 | 59.78 | |
| Llama + NTPLLM=Llama, Training Dataset=MedAlpaca-263K, Training Recipe=NTP, Evaluation Protocol=three-shot setting2024.10 | 59.31 | 54.52 | 75.4 | 53.65 | 60.72 | |
| Mistral + SFTMixLLM=Mistral, Training Dataset=MedAlpaca-263K, Training Recipe=SFTMix, Evaluation Protocol=three-shot setting2024.10 | 51.77 | 45.72 | 77.4 | 49.03 | 55.98 | |
| MistralLLM=Mistral, Training Dataset=MedAlpaca-263K, Evaluation Protocol=three-shot setting2024.10 | 49.18 | 43.94 | 72.33 | 47.98 | 53.36 | |
| Mistral + NTPLLM=Mistral, Training Dataset=MedAlpaca-263K, Training Recipe=NTP, Evaluation Protocol=three-shot setting2024.10 | 49.1 | 44.62 | 75.4 | 48.15 | 54.32 | |
| BioMistralLLM=7B, Evaluation Protocol=three-shot setting2024.10 | 43.86 | 37.58 | 50.13 | 44.14 | 43.93 | |
| MedAlpacaLLM=7B, Evaluation Protocol=three-shot setting2024.10 | 38.94 | 33.96 | 57.2 | 34.9 | 41.25 | |
| BioMedGPTLLM=7B, Evaluation Protocol=three-shot setting2024.10 | 38.62 | 34.72 | 58.27 | 35.57 | 41.8 | |
| MeditronLLM=7B, Evaluation Protocol=three-shot setting2024.10 | 35.09 | 26.73 | 56.93 | 34.03 | 38.2 | |
| PMC-LLaMALLM=7B, Evaluation Protocol=three-shot setting2024.10 | 27.94 | 21.24 | 54.87 | 24.57 | 32.16 |