Question Answering on MedQA
94.8AccuracyPulseMind-72B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PulseMind-72BModel Category=Open-source, Model Scale=~72B2026.01 | 94.8 | — | |
| Gemini 2.5Retrieval Configuration=Our Framework2026.02 | 93.7 | — | |
| InternVL3-78BModel Category=Open-source, Model Scale=~72B2026.01 | 93.3 | — | |
| PulseMind-32BModel Category=Open-source, Model Scale=~32B2026.01 | 92.9 | — | |
| Qwen2.5VL-72BModel Category=Open-source, Model Scale=~72B2026.01 | 91.3 | — | |
| GPT-4.1-miniRetrieval Configuration=Our Framework2026.02 | 91.2 | — | |
| Claude Sonnet 42025.10 | 89.5 | — | |
| Gemini 2.5Retrieval Configuration=with RAG2026.02 | 89.2 | — | |
| GPT-4.12025.10 | 89 | — | |
| GPT-4.1-miniRetrieval Configuration=with RAG2026.02 | 88.5 | — | |
| Gemini 2.5Retrieval Configuration=Baselines without Retrieval2026.02 | 86.7 | — | |
| o1Model Category=Proprietary2026.01 | 86.6 | — | |
| DeepSeek-ChatRetrieval Configuration=Our Framework2026.02 | 86.4 | — | |
| Gemini2.5-proModel Category=Proprietary2026.01 | 85.6 | — | |
| DeepSeek-ChatRetrieval Configuration=with RAG2026.02 | 85.1 | — | |
| GPT-4.1-miniRetrieval Configuration=Baselines without Retrieval2026.02 | 84.6 | — | |
| MedAgentsBackbone=GPT-4, Evaluation Protocol=zero-shot setting2023.11 | 83.7 | — | |
| DOCTOR-R12025.10 | 83.5 | — | |
| Few-shot CoT + SCBackbone=GPT-4, Evaluation Protocol=few-shot setting2023.11 | 82.9 | — | |
| MeerkatModel size=70B2024.03 | 82.6 | — | |
| Baichuan-M2-32B2025.10 | 81.5 | — | |
| GPT-4Shots=5-shot2024.03 | 81.4 | — | |
| DeepSeek-ChatRetrieval Configuration=Baselines without Retrieval2026.02 | 80.9 | — | |
| 5x HumanNumber of Human Annotations=1000, Number of GPT-3.5 Annotations=02023.10 | 80.23 | — | |
| Med42-v2-8B2025.10 | 77.5 | — | |
| Few-shotBackbone=GPT-4, Evaluation Protocol=few-shot setting2023.11 | 76.6 | — | |
| UltraMedical-8B2025.10 | 75 | — | |
| Lingshu-32BModel Category=Open-source, Model Scale=~32B2026.01 | 74.7 | — | |
| Zero-shot CoT + SCBackbone=GPT-4, Evaluation Protocol=zero-shot setting2023.11 | 74.5 | — | |
| MeerkatModel size=8B2024.03 | 74 | — | |
| TARSEBackbone=14B2026.03 | 73.8 | — | |
| InternVL3-38BModel Category=Open-source, Model Scale=~32B2026.01 | 73.5 | — | |
| Few-shot CoTBackbone=GPT-4, Evaluation Protocol=few-shot setting2023.11 | 73.3 | — | |
| Zero-shotBackbone=GPT-4, Evaluation Protocol=zero-shot setting2023.11 | 73 | — | |
| 4x HumanNumber of Human Annotations=800, Number of GPT-3.5 Annotations=02023.10 | 72.49 | — | |
| MedReason-8B2026.02 | 72.4 | — | |
| DRLBackbone=Qwen3-8B, top-k=best-performing2026.02 | 72.2 | — | |
| Qwen2.5VL-32BModel Category=Open-source, Model Scale=~32B2026.01 | 71.6 | — | |
| MeerkatModel size=7B2024.03 | 70.6 | — | |
| Qwen3-8B2026.02 | 70.2 | — | |
| TARSEBackbone=7B2026.03 | 70.1 | — | |
| IMFLNumber of Human Annotations=200, Number of GPT-3.5 Annotations=8002023.10 | 67.98 | — | |
| Few-shot CoT + SCBackbone=Flan-PaLM, Evaluation Protocol=few-shot setting2023.11 | 67.6 | — | |
| MedAgentsBackbone=GPT-3.5, Evaluation Protocol=zero-shot setting2023.11 | 64.1 | — | |
| LLAVA-med-34BModel Category=Open-source, Model Scale=~32B2026.01 | 63.5 | — | |
| Qwen3-8Bbase_model=true2025.10 | 63.5 | — | |
| rStar-Qwen2.5Backbone=14B2026.03 | 63.2 | — | |
| MedPRM-8B2026.02 | 62.2 | — | |
| Few-shot CoT + SCBackbone=GPT-3.5, Evaluation Protocol=few-shot setting2023.11 | 62.1 | — | |
| Zero-shot CoTBackbone=GPT-4, Evaluation Protocol=zero-shot setting2023.11 | 61.8 | — | |
| Gemini-2.5-Flash2025.10 | 61.5 | — | |
| Zero-shot CoT + SCBackbone=GPT-3.5, Evaluation Protocol=zero-shot setting2023.11 | 61.3 | — | |
| HuatuoGPT-o1-8B2026.02 | 60.8 | — | |
| Qwen2.5-InstructBackbone=14B2026.03 | 60.8 | — | |
| Few-shot CoTBackbone=Flan-PaLM, Evaluation Protocol=few-shot setting2023.11 | 60.3 | — | |
| 3x HumanNumber of Human Annotations=600, Number of GPT-3.5 Annotations=02023.10 | 58.31 | — | |
| rStar-Qwen2.5Backbone=7B2026.03 | 58.1 | — | |
| DoctorAgent-RL2025.10 | 58 | — | |
| GPT-3.5 Turbo 1106Evaluation Protocol=Few-shot, Few-shot settings=3-shot2024.02 | 57.71 | — | |
| HuatuoGPT-vision-34BModel Category=Open-source, Model Scale=~32B2026.01 | 57.4 | — | |
| GPT-4oModel Category=Proprietary2026.01 | 55.7 | — | |
| Few-shot CoTBackbone=GPT-3.5, Evaluation Protocol=few-shot setting2023.11 | 55.3 | — | |
| MedAlpaca-7B2023.11 | 55.2 | — | |
| Few-shotBackbone=GPT-3.5, Evaluation Protocol=few-shot setting2023.11 | 54.7 | — | |
| Zero-shotBackbone=GPT-3.5, Evaluation Protocol=zero-shot setting2023.11 | 54.3 | — | |
| BioMistralModel size=7B2024.03 | 54.3 | — | |
| i-MedRAG-Qwen2.5Backbone=7B2026.03 | 54.3 | — | |
| GPT-3.5Shots=5-shot2024.03 | 53.6 | — | |
| DRLBackbone=LLaMA-3.1-8B-Instruct, top-k=best-performing2026.02 | 53.6 | — | |
| Qwen2.5-InstructBackbone=7B2026.03 | 53.2 | — | |
| 2x HumanNumber of Human Annotations=400, Number of GPT-3.5 Annotations=02023.10 | 52.53 | — | |
| LLaMA-3.1-8B-Instruct2026.02 | 51.2 | — | |
| BioMistral 7B DAREEvaluation Protocol=SFT, Few-shot settings=3-shot, Merging Strategy=DARE2024.02 | 51.1 | — | |
| MedRAG-Qwen2.5Backbone=7B2026.03 | 51 | — | |
| BioMistral 7B SLERPEvaluation Protocol=SFT, Few-shot settings=3-shot, Merging Strategy=SLERP2024.02 | 50.8 | — | |
| BioMistral 7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 50.6 | — | |
| BioMistral 7B EnsembleEvaluation Protocol=SFT, Few-shot settings=3-shot, Merging Strategy=Ensemble2024.02 | 50.6 | — | |
| BioMedGPT-10B2023.11 | 50.4 | — | |
| BioMedLM-2.7B2023.11 | 50.3 | — | |
| MediTronModel size=7B2024.03 | 50.2 | — | |
| BioMistral 7B TIESEvaluation Protocol=SFT, Few-shot settings=3-shot, Merging Strategy=TIES2024.02 | 49.5 | — | |
| 1x HumanNumber of Human Annotations=200, Number of GPT-3.5 Annotations=02023.10 | 48.77 | — | |
| All GPT-3.5Number of Human Annotations=0, Number of GPT-3.5 Annotations=10002023.10 | 48.03 | — | |
| Zero-shot CoTBackbone=GPT-3.5, Evaluation Protocol=zero-shot setting2023.11 | 44.3 | — | |
| BioMedGPT-LM-7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 42.5 | — | |
| Mistral 7B InstructEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 42 | — | |
| MediTron-7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 41.6 | — | |
| MedAlpaca 7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 40.1 | — | |
| SciBERT (large)2023.11 | 39.2 | — | |
| BioBERT (large)2023.11 | 36.7 | — | |
| BERT (large)2023.11 | 33.6 | — | |
| Deep EnsembleEnsemble Size=5, Seed=Single2026.02 | 26.8 | 0.007 | |
| PMC-LLaMA 7BEvaluation Protocol=SFT, Few-shot settings=3-shot2024.02 | 25.5 | — | |
| UAT-LITESeed=Single2026.02 | 24.9 | 0.04 | |
| BaselineSeed=Single2026.02 | 24.6 | 0.033 | |
| Temperature ScalingSeed=Single2026.02 | 24.6 | 0.023 |