Medical Reasoning on MedAgentsBench Hard Subsets
0.52MEDQAMEDCOG-META
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| MEDCOG-METAStrategy=Meta-Cognitive Regulation, Memory Usage=In-Distribution (MEDQA, MEDMCQA, MMLU) / Out-of-Distribution (MMLU-PRO, PUBMEDQA)2026.02 | 0.52 | 0.36 | 0.356 | 0.44 | 0.2 | 0.375 | 0.443 | |
| MEDCOG-ALLStrategy=SCoT+KG+Memory, Memory Usage=In-Distribution (MEDQA, MEDMCQA, MMLU) / Out-of-Distribution (MMLU-PRO, PUBMEDQA)2026.02 | 0.5 | 0.32 | 0.288 | 0.36 | 0.19 | 0.332 | 0.182 | |
| AFLOWStrategy=Agent Workflow2026.02 | 0.48 | 0.31 | 0.384 | 0.37 | 0.18 | 0.345 | 0.124 | |
| MULTIPERSONAStrategy=Multi-Persona2026.02 | 0.45 | 0.25 | 0.37 | 0.42 | 0.15 | 0.328 | 0.153 | |
| MEDAGENTSStrategy=Medical Agents2026.02 | 0.43 | 0.3 | 0.288 | 0.08 | 0.15 | 0.25 | -0.036 | |
| SELF-REFINEStrategy=Self-Refine2026.02 | 0.41 | 0.34 | 0.342 | 0.34 | 0.13 | 0.312 | 0.331 | |
| COTStrategy=Chain-of-Thought2026.02 | 0.39 | 0.3 | 0.26 | 0.35 | 0.1 | 0.28 | — | |
| COT-SCStrategy=Chain-of-Thought with Self-Consistency2026.02 | 0.37 | 0.35 | 0.301 | 0.43 | 0.06 | 0.302 | 0.104 | |
| MDAGENTSStrategy=Multi-Dataset Agents2026.02 | 0.36 | 0.22 | 0.247 | 0.08 | 0.11 | 0.203 | -0.169 | |
| MEDPROMPTStrategy=Medical Prompting2026.02 | 0.34 | 0.26 | 0.26 | 0.22 | 0.11 | 0.238 | -3 | |
| ZERO-SHOTStrategy=Zero-shot2026.02 | 0.32 | 0.25 | 0.247 | 0.21 | 0.09 | 0.223 | — | |
| FEW-SHOTStrategy=Few-shot2026.02 | 0.28 | 0.29 | 0.274 | 0.09 | 0.2 | 0.227 | — |