Medical Reasoning on MedDDx (test)
48.2Basic AccuracyMedLA+LLaMA3.1(8B)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| MedLA+LLaMA3.1(8B)Method Category=Logic Based, Backbone=LLaMA3.1, Model Size=8B2025.09 | 48.2 | 43 | 41.7 | 44.3 | |
| CoT-LLaMA3.1(8B)Method Category=LLM with COT, Prompting Strategy=Chain-of-Thought, Model Size=8B, Reference=Meta20242025.09 | 43.9 | 39.3 | 32.2 | 38.5 | |
| LLaMA3.1(8B)[baseline]Method Category=General & Medical LLM, Model Size=8B, Reference=Meta20242025.09 | 43.4 | 36.8 | 30.6 | 36.9 | |
| CoT-LLaMA3(8B)Method Category=LLM with COT, Prompting Strategy=Chain-of-Thought, Model Size=8B, Reference=Meta20242025.09 | 43.4 | 36.8 | 31.3 | 37.2 | |
| LLaMA3(8B)Method Category=General & Medical LLM, Model Size=8B, Reference=Meta20242025.09 | 42.8 | 31.9 | 30.6 | 35.1 | |
| MDAgentsMethod Category=Multi-Agent, Reference=NeurIPS20242025.09 | 42.1 | 37.5 | 33.4 | 37.7 | |
| Mistral(7B)Method Category=General & Medical LLM, Model Size=7B, Reference=Mistral20232025.09 | 41.2 | 35.6 | 37.5 | 38.1 | |
| MedAgentsMethod Category=Multi-Agent, Reference=ACL20242025.09 | 41 | 35.7 | 32.9 | 36.5 | |
| CoT-Mistral(7B)Method Category=LLM with COT, Prompting Strategy=Chain-of-Thought, Model Size=7B, Reference=Mistral20232025.09 | 40.4 | 36.8 | 37.9 | 38.4 | |
| MedAlpaca(7B)Method Category=General & Medical LLM, Model Size=7B, Reference=BHT20232025.09 | 39.9 | 32.5 | 31.1 | 34.5 | |
| MV-LLaMA3.1(8B)Method Category=Multi-Agent, Evaluation Strategy=Majority Voting, Backbone=LLaMA3.1, Model Size=8B2025.09 | 39.6 | 32.8 | 30.2 | 34.2 | |
| CoT-MedAlpaca(7B)Method Category=LLM with COT, Prompting Strategy=Chain-of-Thought, Model Size=7B, Reference=BHT20232025.09 | 39.5 | 32.1 | 31.2 | 34.3 | |
| DyLANMethod Category=Multi-Agent, Reference=COLM20242025.09 | 39.3 | 33.5 | 31.1 | 34.6 | |
| CoT-Llama3-OB(8B)Method Category=LLM with COT, Prompting Strategy=Chain-of-Thought, Model Size=8B, Reference=Saama20242025.09 | 37 | 33 | 32.7 | 34.2 | |
| MedRAG(70B)Method Category=RAG & Based Methods, Prompting Strategy=RAG, Model Size=70B, Reference=Oxon20242025.09 | 36.5 | 34.8 | 32.7 | 34.7 | |
| QAGNNMethod Category=Graph Based, Reference=NAACL20212025.09 | 29.5 | 26.5 | 25.3 | 27.1 | |
| CoT-LLaMA2(7B)Method Category=LLM with COT, Prompting Strategy=Chain-of-Thought, Model Size=7B, Reference=Meta20232025.09 | 28.9 | 26.5 | 22.9 | 26.1 | |
| DragonMethod Category=Graph Based, Reference=NeurIPS20222025.09 | 28.6 | 24.7 | 24 | 25.8 | |
| LLaMA2(13B)Method Category=General & Medical LLM, Model Size=13B, Reference=Meta20232025.09 | 28.6 | 33.8 | 31.7 | 31.4 | |
| CoT-LLaMA2(13B)Method Category=LLM with COT, Prompting Strategy=Chain-of-Thought, Model Size=13B, Reference=Meta20232025.09 | 25.6 | 26.3 | 24.3 | 25.4 | |
| KG-Rank(13B)Method Category=RAG & Based Methods, Prompting Strategy=RAG, Model Size=13B, Reference=ACL-w20242025.09 | 25.3 | 25.6 | 23.4 | 24.8 | |
| Self-RAG(13B)Method Category=RAG & Based Methods, Prompting Strategy=RAG, Model Size=13B, Reference=ICLR20242025.09 | 24.9 | 29 | 26.6 | 26.8 | |
| JointLKMethod Category=Graph Based, Reference=NAACL20222025.09 | 24.7 | 25.3 | 24.4 | 24.8 | |
| Llama3-OB(8B)Method Category=General & Medical LLM, Model Size=8B, Evaluation Strategy=OpenBioLLM, Reference=Saama20242025.09 | 23.8 | 23.5 | 22.9 | 23.4 | |
| Self-RAG(7B)Method Category=RAG & Based Methods, Prompting Strategy=RAG, Model Size=7B, Reference=ICLR20242025.09 | 23.8 | 19.9 | 22.4 | 22 | |
| LLaMA2(7B)Method Category=General & Medical LLM, Model Size=7B, Reference=Meta20232025.09 | 21.5 | 19.8 | 19.2 | 20.2 | |
| CoT-PMC-LLaMA(7B)Method Category=LLM with COT, Prompting Strategy=Chain-of-Thought, Model Size=7B, Reference=SJTU20242025.09 | 8.8 | 7.7 | 6.3 | 7.6 | |
| PMC-LLaMA(7B)Method Category=General & Medical LLM, Model Size=7B, Reference=SJTU20242025.09 | 8.7 | 8.6 | 7.9 | 8.4 |