Medical Diagnosis Accuracy on Public English and In-house Chinese Dataset Combined
48.9Overall Diagnosis AccuracyOurs-SFT-32B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Ours-SFT-32BModel Category=Our Models, Training Strategy=SFT, Parameter Scale=32B2026.07 | 48.9 | 47.6 | |
| QwQ-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 48.2 | 46.1 | |
| M2-32BModel Category=Medical Reasoning Models, Parameter Scale=32B2026.07 | 46 | 41.7 | |
| Qwen2.5-72BModel Category=Chat models, Parameter Scale=72B2026.07 | 45 | 42.1 | |
| Qwen3-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 39.3 | 35.2 | |
| Ours-RL-7BModel Category=Our Models, Training Strategy=RLVR, Parameter Scale=7B2026.07 | 38.2 | 35.3 | |
| DiagAgent-14BModel Category=Medical Reasoning Models, Parameter Scale=14B2026.07 | 36.9 | 33.2 | |
| Qwen2.5-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 31.4 | 29.9 | |
| Huatuo-7BModel Category=Medical Reasoning Models, Parameter Scale=7B2026.07 | 20.7 | 15.9 |