Question Answering on PubMedQA
83.6AccuracyMulti-Agent Medical Decision Consensus Matrix System
Evaluation Results
| Method | Links | |
|---|---|---|
| Multi-Agent Medical Decision Consensus Matrix System2025.12 | 83.6 | |
| URCALLM=GPT-42025.05 | 81.1 | |
| BioGPT-LargeFine-tuned=true, Model Parameters=1.5B, Backbone=GPT-2 XL2022.10 | 81 | |
| GraphRAGLLM=GPT-42025.05 | 80.6 | |
| RankRAGDescription=Fine-tuned RAG pipeline for domain2025.05 | 79.8 | |
| RAGLLM=GPT-42025.05 | 79.6 | |
| TeamMedAgents2025.12 | 79.2 | |
| MDAgents2025.12 | 78.5 | |
| MetaGen BlendedRAGStatus=Non Fine-Tuned, Description=Leverages metadata augmentation for retrieval2025.05 | 77.9 | |
| Weighted VotingMulti-Agent Aggregation Method=Weighted Voting2025.12 | 77.9 | |
| BioMistral 7B SLERPEvaluation Protocol=SFT, Few-shot settings=3-shot, Merging Strategy=SLERP2024.02 | 77.8 | |
| BioMistral 7B DAREEvaluation Protocol=SFT, Few-shot settings=3-shot, Merging Strategy=DARE2024.02 | 77.7 | |
| GalacticaModel Size=120B, Shots=0, Domain=in-domain2022.11 | 77.6 | |
| BioMistral 7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 77.5 | |
| BioMistral 7B EnsembleEvaluation Protocol=SFT, Few-shot settings=3-shot, Merging Strategy=Ensemble2024.02 | 77.5 | |
| BioMistral 7B TIESEvaluation Protocol=SFT, Few-shot settings=3-shot, Merging Strategy=TIES2024.02 | 77.5 | |
| No RAGLLM=GPT-42025.05 | 77.5 | |
| Few-shot CoTBackbone=Flan-PaLM, Evaluation Protocol=few-shot setting2023.11 | 77.2 | |
| ModelMergeBackbone Model=Lingshu-7B2025.12 | 77.2 | |
| MedAgentsBackbone=GPT-4, Evaluation Protocol=zero-shot setting2023.11 | 76.8 | |
| BioMedGPT-LM-7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 76.8 | |
| Majority VotingMulti-Agent Aggregation Method=Majority Voting2025.12 | 76.4 | |
| Zero-shotBackbone=GPT-4, Evaluation Protocol=zero-shot setting2023.11 | 76.2 | |
| Borda CountMulti-Agent Aggregation Method=Borda Count2025.12 | 76.1 | |
| OriginalBackbone Model=Lingshu-7B2025.12 | 76 | |
| Chain-of-ThoughtPrompting Strategy=Chain-of-Thought2025.12 | 75.8 | |
| Zero-shot CoT + SCBackbone=GPT-3.5, Evaluation Protocol=zero-shot setting2023.11 | 75.7 | |
| Mistral 7B InstructEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 75.7 | |
| Few-shot CoT + SCBackbone=GPT-4, Evaluation Protocol=few-shot setting2023.11 | 75.6 | |
| PSIBackbone Model=Lingshu-7B2025.12 | 75.6 | |
| Zero-shot CoT + SCBackbone=GPT-4, Evaluation Protocol=zero-shot setting2023.11 | 75.3 | |
| Few-shot CoT + SCBackbone=Flan-PaLM, Evaluation Protocol=few-shot setting2023.11 | 75.2 | |
| 5x HumanNumber of Human Annotations=1000, Number of GPT-3.5 Annotations=02023.10 | 75.08 | |
| Llama-3.1-8B-InstructSetting=Teacher2026.02 | 75 | |
| Few-shot CoTBackbone=GPT-4, Evaluation Protocol=few-shot setting2023.11 | 74.9 | |
| MediTron-7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 74.9 | |
| Single-AgentEvaluation Protocol=Few-shot2025.12 | 74.3 | |
| ModelMergeBackbone Model=HuatuoGPT-Vision-7B2025.12 | 74.2 | |
| AlzheimerRAGDescription=Multimodal integration (text + visuals)2025.05 | 74 | |
| RESTABackbone Model=Lingshu-7B2025.12 | 73.8 | |
| IMFLNumber of Human Annotations=200, Number of GPT-3.5 Annotations=8002023.10 | 73.76 | |
| Zero-shotBackbone=GPT-3.5, Evaluation Protocol=zero-shot setting2023.11 | 73.7 | |
| BLOOMShots=0, Domain=in-domain2022.11 | 73.6 | |
| MedAlpaca 7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 73.6 | |
| Few-shot CoT + SCBackbone=GPT-3.5, Evaluation Protocol=few-shot setting2023.11 | 73.4 | |
| Few-shotBackbone=GPT-4, Evaluation Protocol=few-shot setting2023.11 | 73.4 | |
| RAFT (LLaMA2-7B)Description=Retrieval-augmented fine-tuning, Base Model=LLaMA2-7B2025.05 | 73.3 | |
| BioMistral-7BSetting=Teacher2026.02 | 73.25 | |
| OriginalBackbone Model=HuatuoGPT-Vision-7B2025.12 | 73 | |
| MedAgentsBackbone=GPT-3.5, Evaluation Protocol=zero-shot setting2023.11 | 72.9 | |
| PMC-LLaMA 7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 72.9 | |
| 4x HumanNumber of Human Annotations=800, Number of GPT-3.5 Annotations=02023.10 | 72.66 | |
| GPT-3.5 Turbo 1106Evaluation Protocol=Few-shot, Few-shot settings=3-shot2024.02 | 72.66 | |
| PSIBackbone Model=HuatuoGPT-Vision-7B2025.12 | 72 | |
| Single-AgentEvaluation Protocol=Zero-shot2025.12 | 72 | |
| DSF + RAGDescription=Domain-specific fine-tuning for RAG2025.05 | 71.6 | |
| FLAN-T5 xlargeSetting=Teacher2026.02 | 71.5 | |
| Few-shot CoTBackbone=GPT-3.5, Evaluation Protocol=few-shot setting2023.11 | 71.4 | |
| Zero-shot CoTBackbone=GPT-4, Evaluation Protocol=zero-shot setting2023.11 | 71 | |
| MEDRAG + GPT-4Description=Corpus-optimized medical retrieval, Base Model=GPT-42025.05 | 70.6 | |
| OPTShots=0, Domain=in-domain2022.11 | 70.2 | |
| 3x HumanNumber of Human Annotations=600, Number of GPT-3.5 Annotations=02023.10 | 70.15 | |
| RESTABackbone Model=HuatuoGPT-Vision-7B2025.12 | 70 | |
| 2x HumanNumber of Human Annotations=400, Number of GPT-3.5 Annotations=02023.10 | 67.64 | |
| Few-shotBackbone=GPT-3.5, Evaluation Protocol=few-shot setting2023.11 | 67.6 | |
| All GPT-3.5Number of Human Annotations=0, Number of GPT-3.5 Annotations=10002023.10 | 66.87 | |
| 1x HumanNumber of Human Annotations=200, Number of GPT-3.5 Annotations=02023.10 | 65.28 | |
| OLTQAsetting=zero-shot2023.05 | 64.4 | |
| OLTQAFull Config=True2023.05 | 64.4 | |
| Teacher SelectionSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 63.5 | |
| PLM ClassifierSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 62.5 | |
| OLTQA (w/o Pm)setting=zero-shot2023.05 | 62.07 | |
| Zero-shot CoTBackbone=GPT-3.5, Evaluation Protocol=zero-shot setting2023.11 | 61.3 | |
| Similarity-based RouterSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 61 | |
| Dense ModelBase Model=Llama3.1-8B, Fine-tuning Protocol=LoRA, Sparsity=0%2026.01 | 60.7 | |
| Dense ModelBase Model=Mistral-7B, Fine-tuning Protocol=LoRA, Sparsity=0%2026.01 | 60.7 | |
| EPRsetting=zero-shot2023.05 | 59.67 | |
| GradPrunerBase Model=Llama3.1-8B, Fine-tuning Protocol=LoRA, Sparsity=40%2026.01 | 59.4 | |
| Dense ModelBase Model=Llama3.1-8B, Fine-tuning Protocol=FFT, Sparsity=0%2026.01 | 59.3 | |
| FT(Llama3.2)Base Model=Llama3.2-3B, Fine-tuning Protocol=LoRA, Sparsity=0%2026.01 | 59.2 | |
| FT(Llama3.2)Base Model=Llama3.2-3B, Fine-tuning Protocol=FFT, Sparsity=0%2026.01 | 59.1 | |
| GradPrunerBase Model=Llama3.1-8B, Fine-tuning Protocol=FFT, Sparsity=40%2026.01 | 59.1 | |
| Dense ModelBase Model=Mistral-7B, Fine-tuning Protocol=FFT, Sparsity=0%2026.01 | 59.1 | |
| GradPrunerBase Model=Mistral-7B, Fine-tuning Protocol=LoRA, Sparsity=40%2026.01 | 58.8 | |
| GradPrunerBase Model=Mistral-7B, Fine-tuning Protocol=FFT, Sparsity=40%2026.01 | 58.6 | |
| ProQAsetting=zero-shot2023.05 | 58 | |
| TTARAGBackbone=Llama-3.1-8b-it2026.01 | 57.4 | |
| Muppetsetting=zero-shot2023.05 | 56.73 | |
| SATBase Model=Llama3.1-8B, Fine-tuning Protocol=FFT, Sparsity=40%2026.01 | 56.7 | |
| Plackett-Luce RankingSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 56.5 | |
| OLTQA (w/o Pk)setting=zero-shot2023.05 | 56.27 | |
| APTBase Model=Llama3.1-8B, Fine-tuning Protocol=LoRA, Sparsity=40%2026.01 | 56.1 | |
| SATBase Model=Mistral-7B, Fine-tuning Protocol=FFT, Sparsity=40%2026.01 | 56.1 | |
| LLMPrunerBase Model=Llama3.1-8B, Fine-tuning Protocol=FFT, Sparsity=40%2026.01 | 56 | |
| LacoBase Model=Llama3.1-8B, Fine-tuning Protocol=FFT, Sparsity=40%2026.01 | 55.6 | |
| MINITRONBase Model=Llama3.1-8B, Fine-tuning Protocol=FFT, Sparsity=40%2026.01 | 55.5 | |
| APTBase Model=Mistral-7B, Fine-tuning Protocol=LoRA, Sparsity=40%2026.01 | 55.5 | |
| LacoBase Model=Llama3.1-8B, Fine-tuning Protocol=LoRA, Sparsity=40%2026.01 | 55.3 | |
| LacoBase Model=Mistral-7B, Fine-tuning Protocol=FFT, Sparsity=40%2026.01 | 55.2 | |
| PLM ClassifierSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 55 |