Medical Question Answering on MedQA
93.88AccuracyToolTree
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ToolTreeBackbone Model=GPT-4o2026.03 | 93.88 | — | |
| OctoToolsBackbone Model=GPT-4o2026.03 | 92.17 | — | |
| Multi-Agent Medical Decision Consensus Matrix System2025.12 | 91.7 | — | |
| ToolTreeBackbone Model=GPT-4o-mini2026.03 | 91.13 | — | |
| CoT PromptingModel=GPT-4o, Configuration=Chain-of-Thought Prompting2026.03 | 89.75 | — | |
| GPT-5.12026.01 | 89.55 | — | |
| Base Model OnlyModel=GPT-4o, Configuration=No Augmentation2026.03 | 89 | — | |
| HuatuoGPT-o1-72BParameter Scale=72B2025.04 | 88.85 | — | |
| TeamMedAgents2025.12 | 88.1 | — | |
| Naive RAGModel=GPT-4o, Configuration=Standard Retrieval Augmented Generation2026.03 | 87.75 | — | |
| MDAgents2025.12 | 87.3 | — | |
| HuatuoGPT-o1-70BParameter Scale=70B2025.04 | 86.8 | — | |
| HuggingGPTBackbone Model=GPT-4o2026.03 | 86.73 | — | |
| OctoToolsBackbone Model=GPT-4o-mini2026.03 | 86.18 | — | |
| Weighted VotingMulti-Agent Aggregation Method=Weighted Voting2025.12 | 86.1 | — | |
| GroupRAGModel=GPT-4o, Configuration=Group-aware Retrieval and Reasoning2026.03 | 85.25 | — | |
| Majority VotingMulti-Agent Aggregation Method=Majority Voting2025.12 | 85.2 | — | |
| Borda CountMulti-Agent Aggregation Method=Borda Count2025.12 | 84.8 | — | |
| HuggingGPTBackbone Model=GPT-4o-mini2026.03 | 84.33 | — | |
| UltraMedical-70B-3Parameter Scale=70B2025.04 | 83.9 | — | |
| Chain-of-ThoughtPrompting Strategy=Chain-of-Thought2025.12 | 83.5 | — | |
| m1-32B-1KParameter Scale=32B, Thinking Budget=1K2025.04 | 83.5 | — | |
| Few-ShotBackbone Model=GPT-4o2026.03 | 83.2 | — | |
| Primitives-based MASModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 82.7 | — | |
| Single-AgentEvaluation Protocol=Few-shot2025.12 | 82.1 | — | |
| Primitives-based MASModel=DeepSeek-R1-Distill Llama-70B2026.02 | 81.9 | — | |
| LatentMASModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 81.2 | — | |
| Single-AgentEvaluation Protocol=Zero-shot2025.12 | 80.2 | — | |
| SGR-LLaMA-3.3-70BBackbone=LLaMA-3.3-70B2026.01 | 79.81 | — | |
| TextMASModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 79.6 | — | |
| VotingModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 79.3 | — | |
| Few-ShotBackbone Model=GPT-4o-mini2026.03 | 79.14 | — | |
| TextMASModel=DeepSeek-R1-Distill Llama-70B2026.02 | 77.8 | — | |
| VotingModel=DeepSeek-R1-Distill Llama-70B2026.02 | 77.8 | — | |
| MA-RAG-extBackbone=Qwen3-8B, Ranking Agent=extrinsic verification2026.02 | 77.1 | — | |
| MA-RAG-intBackbone=Qwen3-8B, Ranking Agent=intrinsic uncertainty2026.02 | 77 | — | |
| Primitives-based MASModel=Qwen3-8B2026.02 | 76.7 | — | |
| Qwen2.5-72B-InstructParameter Scale=72B, Chain-of-Thought (CoT) Prompting=true2025.04 | 76.43 | — | |
| m1-7B-23KParameter Scale=7B, Thinking Budget=23K2025.04 | 75.81 | — | |
| Qwen3-8B-thinkingModel Parameters=8B, Reasoning Mode=Thinking2026.01 | 75.8 | — | |
| UltraMedical-8B-3.1Parameter Scale=8B2025.04 | 75.73 | — | |
| LatentMASModel=Qwen3-8B2026.02 | 75.3 | — | |
| Qwen2.5-32B-InstructParameter Scale=32B, Chain-of-Thought (CoT) Prompting=false2025.04 | 75.26 | — | |
| OpenBioLLM-70BParameter Scale=70B2025.04 | 75.1 | — | |
| TextMASModel=Qwen3-8B2026.02 | 75 | — | |
| PlanningModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 75 | — | |
| Qwen2.5-32B-InstructParameter Scale=32B, Chain-of-Thought (CoT) Prompting=true2025.04 | 74.86 | — | |
| HuatuoGPT-o1-8BParameter Scale=8B2025.04 | 74.78 | — | |
| Multi-RefineBackbone=Qwen3-8B2026.02 | 74.7 | — | |
| Qwen2.5-72B-InstructParameter Scale=72B, Chain-of-Thought (CoT) Prompting=false2025.04 | 74.55 | — | |
| PlanningModel=DeepSeek-R1-Distill Llama-70B2026.02 | 74 | — | |
| SCBackbone=Qwen3-8B2026.02 | 73.3 | — | |
| ReviewModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 73 | — | |
| ReviewModel=DeepSeek-R1-Distill Llama-70B2026.02 | 72.9 | — | |
| FLAREBackbone=Qwen3-8B2026.02 | 72.7 | — | |
| GroupRAGModel=L3.1-8B(trained), Configuration=Group-aware Retrieval and Reasoning2026.03 | 71.75 | — | |
| HuatuoGPT-o1-7BParameter Scale=7B2025.04 | 71.56 | — | |
| Qwen3-8BBackbone=Qwen3-8B2026.02 | 71.1 | — | |
| HuatuoGPT-o1-8BBackbone=HuatuoGPT-o1-8B2026.02 | 71.1 | — | |
| UltraMedical-8B-3Parameter Scale=8B2025.04 | 71.09 | — | |
| m1-7B-1KParameter Scale=7B, Thinking Budget=1K2025.04 | 71.01 | — | |
| VotingModel=Qwen3-8B2026.02 | 70.3 | — | |
| TC-RAGBackbone=Qwen3-8B2026.02 | 70 | — | |
| SR-RAGBackbone=Qwen3-8B2026.02 | 69.9 | — | |
| UltraMedical-3.1-8BBackbone=UltraMedical-3.1-8B2026.02 | 69.6 | — | |
| FL-RAGBackbone=Qwen3-8B2026.02 | 69.6 | — | |
| CoTBackbone=Qwen3-8B2026.02 | 69.3 | — | |
| LatentMASModel=DeepSeek-R1-Distill Llama-70B2026.02 | 68.6 | — | |
| Llama-3.1-8BBackbone=Llama-3.1-8B2026.02 | 68.1 | — | |
| SingleModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 68 | — | |
| FS-RAGBackbone=Qwen3-8B2026.02 | 68 | — | |
| DyPRAGBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 67.87 | — | |
| ReFilterBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 67.79 | — | |
| S-RAGBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 67.64 | — | |
| PRAGBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 67.01 | — | |
| G-DesignerAttack Rate=< 50%2026.02 | 67 | 803 | |
| PlanningModel=Qwen3-8B2026.02 | 67 | — | |
| Qwen3-8BModel Parameters=8B2026.01 | 66.8 | — | |
| AgentPruneAttack Rate=< 50%2026.02 | 66.65 | 520 | |
| VanillaBackbone=Qwen2.5-14B-Instruct, Zero-shot evaluation=true2026.02 | 66.22 | — | |
| SingleModel=DeepSeek-R1-Distill Llama-70B2026.02 | 64.5 | — | |
| Qwen2.5-7B-InstructParameter Scale=7B, Chain-of-Thought (CoT) Prompting=true2025.04 | 64.49 | — | |
| ReviewModel=Qwen3-8B2026.02 | 64.2 | — | |
| FullModel=Mixtral-8x7B, Expert Sparsity=0%2025.12 | 62.37 | — | |
| Qwen2.5-7B-InstructParameter Scale=7B, Chain-of-Thought (CoT) Prompting=false2025.04 | 61.51 | — | |
| CoT PromptingModel=L3.1-8B(trained), Configuration=Chain-of-Thought Prompting2026.03 | 61.5 | — | |
| TodyCommAttack Rate=< 50%2026.02 | 61 | 560 | |
| GroupRAGModel=L3.1-8B(base), Configuration=Group-aware Retrieval and Reasoning2026.03 | 61 | — | |
| FullBackbone=Mixtral-8x7B-Instruct, Expert Sparsity=0%2025.12 | 60.88 | — | |
| FullModel=Mixtral-8x7B-Instruct, Expert Sparsity=0%2025.12 | 60.88 | — | |
| MMed-8B-EnInsParameter Scale=8B2025.04 | 60.33 | — | |
| TodyCommAttack Rate== 50%2026.02 | 60.17 | 565 | |
| Med42-8BParameter Scale=8B2025.04 | 59.78 | — | |
| Complete GraphAttack Rate=< 50%2026.02 | 59.67 | 976 | |
| MedLlama3-8B-v2Parameter Scale=8B2025.04 | 59.39 | — | |
| Random GraphAttack Rate=< 50%2026.02 | 58.33 | 764 | |
| Naive RAGModel=L3.1-8B(trained), Configuration=Standard Retrieval Augmented Generation2026.03 | 58.25 | — | |
| TodyCommAttack Rate=> 50%2026.02 | 57.5 | 566 | |
| MMedS-8BParameter Scale=8B2025.04 | 57.19 | — | |
| OpenBioLLM-8BParameter Scale=8B2025.04 | 55.3 | — |