Question Answering on MedQA (test)
89.55AccuracyGPT-5.1
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5.1Model Category=Proprietary LLM2026.01 | 89.55 | |
| GPT-4oModel Category=Proprietary LLM2026.01 | 88.29 | |
| MEDAGENTSBase Model=GPT-42023.11 | 83.7 | |
| SGR-Llama3.3-70BMethod=Self-Graph Reasoning2026.01 | 78.81 | |
| KnowGPTCategory=Ours2023.12 | 78.1 | |
| Claude-3.5-HaikuModel Category=Proprietary LLM2026.01 | 76.36 | |
| GPT-4Category=LLM + Zero-shot2023.12 | 76.3 | |
| MindmapCategory=LLM + KG Prompting2023.12 | 75.1 | |
| Qwen2.5-72BModel Category=Open-source LLM2026.01 | 74.42 | |
| RoGCategory=LLM + KG Prompting2023.12 | 72.6 | |
| CoKCategory=LLM + KG Prompting2023.12 | 72.2 | |
| Llama3Category=LLM + Zero-shot, Parameters=8b2023.12 | 69.7 | |
| DSPy (BootstrapFewShotWithRandomSearch)Solver Model=GPT-3.5-Turbo, Evaluation Setting=few-shot2024.06 | 68.5 | |
| MEDAGENTSBase Model=GPT-3.52023.11 | 64.1 | |
| LLaMA-3.3-70BModel Category=Open-source LLM2026.01 | 63.55 | |
| DSPy (MIPRO v2)Solver Model=GPT-3.5-Turbo, Evaluation Setting=few-shot2024.06 | 62.9 | |
| DSPy (MIPRO v2)Solver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 61.9 | |
| Qwen2.5-7BModel Category=Open-source LLM2026.01 | 59.54 | |
| GPT-3.5 Turbo 1106shot=3-shot2024.02 | 57.71 | |
| UNIPROMPTSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot, Prompt Initialization=Task Description, Search Strategy=Beam2024.06 | 57.1 | |
| UNIPROMPTSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot, Prompt Initialization=Task Description, Search Strategy=Greedy2024.06 | 55.5 | |
| LLaMA-3.1-8BModel Category=Open-source LLM2026.01 | 55.22 | |
| MedAlpacaModel Scale=7B2023.11 | 55.2 | |
| OPROSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 53.3 | |
| Expert PromptSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 53.1 | |
| ProTeGiSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 52.9 | |
| EvokeSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 52.8 | |
| Task DescriptionSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 52.7 | |
| Llama PromptSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot, Reference=Section 3.22024.06 | 52.6 | |
| TextGradSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 50.6 | |
| BioMedGPTModel Scale=10B2023.11 | 50.4 | |
| BioMedLMModel Scale=2.7B2023.11 | 50.3 | |
| CoTSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 50.3 | |
| EvoPromptSolver Model=GPT-3.5-Turbo, Evaluation Setting=zero-shot2024.06 | 50.3 | |
| LLaMA-3.2-3BModel Category=Open-source LLM2026.01 | 49.57 | |
| GPT-3.5Category=LLM + Zero-shot, Implementation=gpt-3.5-turbo2023.12 | 48.7 | |
| BioMistral 7B DAREshot=3-shot, merging_strategy=DARE2024.02 | 47 | |
| BioMistral 7B SLERPshot=3-shot, merging_strategy=SLERP2024.02 | 46.6 | |
| wICIBackbone=Llama-3.1-8B, Training Data Percentage=30%, Data Selection Strategy=In-Context Influence2026.04 | 46.03 | |
| BioMistral 7B Ensembleshot=3-shot, merging_strategy=Ensemble2024.02 | 44.7 | |
| BioMistral 7Bshot=3-shot, options=42024.02 | 44.4 | |
| FullBackbone=Llama-3.1-8B, Training Data Percentage=100%, Data Selection Strategy=All data2026.04 | 44.22 | |
| BioMistral 7B TIESshot=3-shot, merging_strategy=TIES2024.02 | 44 | |
| Mistral 7B Instructshot=3-shot, options=42024.02 | 42.3 | |
| ChatGLM2Category=LLM + Zero-shot2023.12 | 42.2 | |
| JointLKCategory=KG-enhanced LM2023.12 | 40.3 | |
| wICIBackbone=Mistral-7B-v0.3, Training Data Percentage=30%, Data Selection Strategy=In-Context Influence2026.04 | 39.54 | |
| GrapeQACategory=KG-enhanced LM2023.12 | 39.5 | |
| BioMedGPT-LM-7Bshot=3-shot2024.02 | 39.3 | |
| HamQACategory=KG-enhanced LM2023.12 | 38.5 | |
| GreaseLMCategory=KG-enhanced LM2023.12 | 38.5 | |
| QA-GNNCategory=KG-enhanced LM2023.12 | 38.1 | |
| RandomBackbone=Mistral-7B-v0.3, Training Data Percentage=30%, Data Selection Strategy=Random sampling2026.04 | 37 | |
| Bert-largeCategory=LM + Fine tuning2023.12 | 36.7 | |
| BioBERTModel Scale=large2023.11 | 36.7 | |
| ChatGLMCategory=LLM + Zero-shot2023.12 | 36.6 | |
| RoBerta-largeCategory=LM + Fine tuning2023.12 | 36.1 | |
| MedAlpaca 7Bshot=3-shot2024.02 | 35.4 | |
| InternLMCategory=LLM + Zero-shot2023.12 | 34.8 | |
| MediTron-7Bshot=3-shot2024.02 | 34.8 | |
| Bert-baseCategory=LM + Fine tuning2023.12 | 34.4 | |
| FullBackbone=Mistral-7B-v0.3, Training Data Percentage=100%, Data Selection Strategy=All data2026.04 | 34.32 | |
| Llama2Category=LLM + Zero-shot, Parameters=7b2023.12 | 34 | |
| BaichuanCategory=LLM + Zero-shot, Parameters=7B2023.12 | 31.9 | |
| RandomBackbone=Llama-3.1-8B, Training Data Percentage=30%, Data Selection Strategy=Random sampling2026.04 | 29.53 | |
| GPT-3Category=LLM + Zero-shot, Implementation=text-davinci-0022023.12 | 28.9 | |
| PMC-LLAMA 7Bshot=3-shot2024.02 | 27.6 |