Clinical Reasoning on MIMIC-CDM-FI
90.1AccuracyDeepSeek-V4
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-V42026.05 | 90.1 | |
| MedGuideX-9BBackbone=Qwen3.5-9B2026.05 | 85.2 | |
| GPT-5.02026.05 | 85.1 | |
| RL with CPG-Derived Process RewardsBackbone=Qwen3.5-9B2026.05 | 83.4 | |
| Qwen3.5-9BPrompting Strategy=RAG with Guidelines2026.05 | 82.5 | |
| MedGuideX-4BBackbone=Qwen3.5-4B2026.05 | 82.4 | |
| Qwen3.5-4BPrompting Strategy=RAG with Guidelines2026.05 | 81.9 | |
| Claude-Haiku-4.52026.05 | 81.6 | |
| Qwen3.5-9BPrompting Strategy=Zero-shot2026.05 | 81.6 | |
| Qwen3.5-4BPrompting Strategy=3-Shot In-Context Learning2026.05 | 81.4 | |
| Llama-Aloe-Beta-8B2026.05 | 80.8 | |
| Qwen3.5-4BPrompting Strategy=Zero-shot2026.05 | 80.4 | |
| MedGemma-1.5-4B2026.05 | 80.1 | |
| Qwen3.5-9BPrompting Strategy=3-Shot In-Context Learning2026.05 | 80.1 | |
| Hulu-Med-7B2026.05 | 78.7 | |
| Fine-tuning with CPGBackbone=Qwen3.5-9B2026.05 | 78.4 | |
| Lingshu-7B2026.05 | 78 | |
| Kimi-2.62026.05 | 74.6 | |
| MedReason-8B2026.05 | 73 | |
| MediPhi-Instruct-4B2026.05 | 72.6 | |
| HuatuoGPT-o1-8B2026.05 | 68.1 | |
| CPGPromptBackbone=Qwen3.5-9B2026.05 | 67.6 | |
| Clinical-R1-3B2026.05 | 62.3 | |
| Llava-Med-v1.5-7B2026.05 | 60.1 | |
| MedAlpaca-7B2026.05 | 41 | |
| BioMistral-7B2026.05 | 35.6 |