Multi-choice Medical QA on Multi-choice medical QA benchmarks (test)
70.7MMLU-Med AccuracyMedLA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| MedLACategory=Logic Based, Backbone=LLaMA3.1(8B)2025.09 | 70.7 | 62.6 | 76.5 | 69.9 | |
| LLaMA3.1Category=General & Medical LLMs with COT, Chain of Thought (COT)=true, Model Scale=8B2025.09 | 68.1 | 54.9 | 70.6 | 64.5 | |
| LLaMA3.1 (baseline)Category=General & Medical LLMs, Model Scale=8B2025.09 | 67.7 | 56.3 | 68.7 | 64.2 | |
| LLaMA3Category=General & Medical LLMs with COT, Chain of Thought (COT)=true, Model Scale=8B2025.09 | 65.1 | 55.2 | 64.2 | 61.5 | |
| MDAgentsCategory=Multi Agents Methods2025.09 | 65 | 53.4 | 64 | 60.8 | |
| MedAgentsCategory=Multi Agents Methods2025.09 | 64.3 | 53.2 | 64.1 | 60.5 | |
| LLaMA3-OBCategory=General & Medical LLMs, Model Scale=8B2025.09 | 63.6 | 38.3 | 62.2 | 54.7 | |
| MistralCategory=General & Medical LLMs, Model Scale=7B2025.09 | 63.4 | 47.7 | 64.4 | 58.5 | |
| LLaMA3Category=General & Medical LLMs, Model Scale=8B2025.09 | 63.4 | 56.6 | 65.4 | 61.8 | |
| MistralCategory=General & Medical LLMs with COT, Chain of Thought (COT)=true, Model Scale=7B2025.09 | 63.4 | 47.4 | 65.1 | 58.6 | |
| DyLANCategory=Multi Agents Methods2025.09 | 62.5 | 51.6 | 63.8 | 59.3 | |
| MedAlpacaCategory=General & Medical LLMs with COT, Chain of Thought (COT)=true, Model Scale=7B2025.09 | 60.3 | 39.9 | 48.5 | 49.6 | |
| MV-LLaMA3.1Category=Multi Agents Methods, Model Scale=8B2025.09 | 60.2 | 46.8 | 65.2 | 57.4 | |
| MedAlpacaCategory=General & Medical LLMs, Model Scale=7B2025.09 | 60 | 40.1 | 49.3 | 49.8 | |
| MedRAGCategory=RAG Based Methods, Model Scale=70B2025.09 | 57.9 | 48.7 | 71.9 | 59.5 | |
| Llama3-OBCategory=General & Medical LLMs with COT, Chain of Thought (COT)=true, Model Scale=8B2025.09 | 57.1 | 39.5 | 64.6 | 53.7 | |
| Self-RAGCategory=RAG Based Methods, Model Scale=13B2025.09 | 50.2 | 40.8 | 64.6 | 51.9 | |
| KG-RankCategory=RAG Based Methods, Model Scale=13B2025.09 | 45.2 | 36.2 | 50.3 | 43.9 | |
| LLaMA2Category=General & Medical LLMs, Model Scale=13B2025.09 | 44.2 | 25.3 | 34.6 | 34.7 | |
| LLaMA2Category=General & Medical LLMs with COT, Chain of Thought (COT)=true, Model Scale=13B2025.09 | 41.5 | 35.4 | 36.3 | 37.7 | |
| LLaMA2Category=General & Medical LLMs, Model Scale=7B2025.09 | 37.6 | 28.1 | 56.8 | 40.8 | |
| Self-RAGCategory=RAG Based Methods, Model Scale=7B2025.09 | 32.2 | 38 | 59.4 | 43.2 | |
| DragonCategory=Graph Based Methods2025.09 | 31.9 | 47.5 | 70.6 | 50 | |
| LLaMA2Category=General & Medical LLMs with COT, Chain of Thought (COT)=true, Model Scale=7B2025.09 | 31.8 | 25.1 | 54.7 | 37.2 | |
| QAGNNCategory=Graph Based Methods2025.09 | 31.7 | 47 | 70.7 | 49.8 | |
| JointLKCategory=Graph Based Methods2025.09 | 28.8 | 42.5 | 70.6 | 47.3 | |
| PMC-LLaMACategory=General & Medical LLMs, Model Scale=7B2025.09 | 20.7 | 24.7 | 34.6 | 26.7 | |
| PMC-LLaMACategory=General & Medical LLMs with COT, Chain of Thought (COT)=true, Model Scale=7B2025.09 | 20.4 | 20.8 | 20.8 | 20.7 |