Medical Question Answering on PubMedQA (Accuracy)
81.4AccuracyHuatuoGPT-o1-70B
Evaluation Results
| Method | Links | |
|---|---|---|
| HuatuoGPT-o1-70BParameter Scale=70B2025.04 | 81.4 | |
| HuatuoGPT-o1-8BParameter Scale=8B2025.04 | 80.1 | |
| UltraMedical-70B-3Parameter Scale=70B2025.04 | 80 | |
| o3Size category=Large Models2025.07 | 80 | |
| HuatuoGPT-o1-72BParameter Scale=72B2025.04 | 79.9 | |
| OpenBioLLM-70BParameter Scale=70B2025.04 | 79.3 | |
| UltraMedical-8B-3.1Parameter Scale=8B2025.04 | 79.2 | |
| OpenBioLLM 70BSize category=Large Models2025.07 | 79 | |
| RewriteBackbone=Llama-3.1-8B, Method Type=Model-based conversion2026.03 | 78.8 | |
| HuatuoGPT-o1-7BParameter Scale=7B2025.04 | 78.6 | |
| FilterBackbone=Llama-3.1-8B, Method Type=Filtering non-convertible questions2026.03 | 78.5 | |
| GPT-4oSize category=Large Models2025.07 | 78.4 | |
| DirectBackbone=Llama-3.1-8B, Method Type=Direct conversion to short-answer2026.03 | 78.3 | |
| Med42-70BParameter Scale=70B2025.04 | 78.1 | |
| BaseBackbone=Llama-3.1-8B, Method Type=Base instruct model2026.03 | 78 | |
| Orig.Backbone=Llama-3.1-8B, Method Type=Standard RLVR on original MCQs2026.03 | 77.8 | |
| BioMistral DARE 7BSize category=Small Models2025.07 | 77.7 | |
| m1-32B-1KParameter Scale=32B, Thinking Budget=1K2025.04 | 77.6 | |
| IDCBackbone=Llama-3.1-8B, Method Type=Iterative Distractor Curation2026.03 | 77.6 | |
| MMedS-8BParameter Scale=8B2025.04 | 77.5 | |
| m1-7B-1KParameter Scale=7B, Thinking Budget=1K2025.04 | 77.5 | |
| DeepSeek R1Size category=Large Models2025.07 | 77.2 | |
| MedGemma 27BSize category=Small Models, test-time scaling=true2025.07 | 76.8 | |
| IQVIA Med-R1 8BSize category=Small Models2025.07 | 76.4 | |
| Gemini 2.5 FlashSize category=Large Models2025.07 | 76.2 | |
| Qwen3-8B + GRPOBackbone=Qwen3-8B, Training=GRPO2026.06 | 76.2 | |
| Med42-8BParameter Scale=8B2025.04 | 76 | |
| m1-7B-23KParameter Scale=7B, Thinking Budget=23K2025.04 | 75.8 | |
| Gemini 2.5 ProSize category=Large Models2025.07 | 75.8 | |
| MEDIC-ADSize=7B2026.03 | 75.6 | |
| MedLlama3-8B-v2Parameter Scale=8B2025.04 | 75.5 | |
| LingshuSize=7B2026.03 | 75.4 | |
| Qwen3-8B + RCSDBackbone=Qwen3-8B, Training=RCSD2026.06 | 75.1 | |
| Citrus-VSize=8B2026.03 | 74.8 | |
| Qwen3-8B + OPSDBackbone=Qwen3-8B, Training=OPSD2026.06 | 74.4 | |
| JSL-MedLlama 3 8B v2.0Size category=Small Models2025.07 | 74.2 | |
| Qwen3-8BBackbone=Qwen3-8B2026.06 | 74.2 | |
| Qwen3-8B + GRPO-RubricsBackbone=Qwen3-8B, Training=GRPO-Rubrics2026.06 | 74.2 | |
| OpenBioLLM 8BSize category=Small Models2025.07 | 74.1 | |
| IDCBackbone=Qwen2-7B, Method Type=Iterative Distractor Curation2026.03 | 74 | |
| MedGemma 4BSize category=Small Models2025.07 | 73.4 | |
| Gemma 3 27BSize category=Small Models2025.07 | 73.4 | |
| FilterBackbone=Qwen2-7B, Method Type=Filtering non-convertible questions2026.03 | 72.9 | |
| Qwen2.5-7B-InstructParameter Scale=7B, Chain-of-Thought (CoT) Prompting=true2025.04 | 72.6 | |
| Orig.Backbone=Qwen2-7B, Method Type=Standard RLVR on original MCQs2026.03 | 72.5 | |
| DirectBackbone=Qwen2-7B, Method Type=Direct conversion to short-answer2026.03 | 72 | |
| Qwen2.5-7B-InstructParameter Scale=7B, Chain-of-Thought (CoT) Prompting=false2025.04 | 71.3 | |
| Qwen2.5-72B-InstructParameter Scale=72B, Chain-of-Thought (CoT) Prompting=true2025.04 | 71.3 | |
| RewriteBackbone=Qwen2-7B, Method Type=Model-based conversion2026.03 | 71.2 | |
| UltraMedical-8B-3Parameter Scale=8B2025.04 | 71.1 | |
| Qwen2.5-72B-InstructParameter Scale=72B, Chain-of-Thought (CoT) Prompting=false2025.04 | 70.8 | |
| OpenBioLLM-8BParameter Scale=8B2025.04 | 70.1 | |
| BaseBackbone=Qwen2-7B, Method Type=Base instruct model2026.03 | 69.5 | |
| ReConcileBase Model=Gemma-3-4B2025.08 | 69.3 | |
| Qwen2.5-32B-InstructParameter Scale=32B, Chain-of-Thought (CoT) Prompting=true2025.04 | 68.9 | |
| TMA-AllComponBase Model=Gemma-3-4B2025.08 | 68.7 | |
| Gemma 3 4BSize category=Small Models2025.07 | 68.4 | |
| Qwen2.5-32B-InstructParameter Scale=32B, Chain-of-Thought (CoT) Prompting=false2025.04 | 68 | |
| Single-Agent BestBase Model=Gemma-3-4B2025.08 | 66.6 | |
| MMed-8B-EnInsParameter Scale=8B2025.04 | 63.8 | |
| MMed-8BParameter Scale=8B2025.04 | 63.4 | |
| ORMModel Size=8B2026.03 | 62 | |
| OVM2026.03 | 59.4 | |
| SEMA-RAGBackbone=deepseek-v3.12026.05 | 59.2 | |
| SEMA-RAGBackbone=gemini-2.0-flash2026.05 | 59.2 | |
| MCNIGModel Size=8B2026.03 | 57.7 | |
| ReFilterBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 56.8 | |
| DyLanBase Model=Gemma-3-4B2025.08 | 56.7 | |
| MedAgentsBase Model=Gemma-3-4B2025.08 | 56.1 | |
| ReFilterBackbone=LLaMA-3-8B-Instruct, Zero-shot evaluation=true2026.02 | 56 | |
| SEMA-RAGBackbone=qwen3-coder-plus2026.05 | 56 | |
| SEMA-RAGBackbone=kimi-k22026.05 | 55.8 | |
| ReFilterBackbone=Qwen2.5-1.5B-Instruct, Zero-shot evaluation=true2026.02 | 55.2 | |
| S-RAGBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 55.2 | |
| DyPRAGBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 55.2 | |
| DyPRAGBackbone=LLaMA-3-8B-Instruct, Zero-shot evaluation=true2026.02 | 54.8 | |
| i-MedRAGBackbone=kimi-k22026.05 | 54.6 | |
| S-RAGBackbone=LLaMA-3-8B-Instruct, Zero-shot evaluation=true2026.02 | 54.2 | |
| QwenPRMModel Size=7B2026.03 | 54.1 | |
| PRAGBackbone=LLaMA-3-8B-Instruct, Zero-shot evaluation=true2026.02 | 54 | |
| VanillaBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 54 | |
| PRAGBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 54 | |
| CoTBackbone=kimi-k22026.05 | 53.6 | |
| MedLlama3-8B-v1Parameter Scale=8B2025.04 | 52.7 | |
| MedRAGBackbone=kimi-k22026.05 | 52.6 | |
| VanillaBackbone=LLaMA-3-8B-Instruct, Zero-shot evaluation=true2026.02 | 51.8 | |
| PRAGBackbone=Qwen2.5-1.5B-Instruct, Zero-shot evaluation=true2026.02 | 51.8 | |
| S-RAGBackbone=Qwen2.5-1.5B-Instruct, Zero-shot evaluation=true2026.02 | 51.6 | |
| i-MedRAGBackbone=gemini-2.0-flash2026.05 | 51.2 | |
| SEMA-RAGBackbone=glm-4.0-flash2026.05 | 51.2 | |
| i-MedRAGBackbone=deepseek-v3.12026.05 | 50.6 | |
| MedCPTBackbone=kimi-k22026.05 | 50.2 | |
| DyPRAGBackbone=Qwen2.5-1.5B-Instruct, Zero-shot evaluation=true2026.02 | 50 | |
| VanillaBackbone=Qwen2.5-1.5B-Instruct, Zero-shot evaluation=true2026.02 | 49.8 | |
| Granite PRM v22026.03 | 49.8 | |
| MathShepherd2026.03 | 49.7 | |
| MedRAGBackbone=qwen3-coder-plus2026.05 | 49.2 | |
| IGModel Size=8B2026.03 | 49 | |
| MDAgentsBase Model=Gemma-3-4B2025.08 | 48.7 | |
| i-MedRAGBackbone=qwen3-coder-plus2026.05 | 48.6 |