Medical Diagnosis Accuracy on Public English Dataset Follow-up Consultation
44.5R1QwQ-32B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| QwQ-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 44.5 | 42.7 | 40.9 | 42.7 | |
| Ours-SFT-32BModel Category=Our Models, Training Strategy=SFT, Parameter Scale=32B2026.07 | 43.6 | 45.5 | 39.1 | 42.7 | |
| Qwen2.5-72BModel Category=Chat models, Parameter Scale=72B2026.07 | 41.8 | 50.9 | 39.1 | 43.9 | |
| M2-32BModel Category=Medical Reasoning Models, Parameter Scale=32B2026.07 | 40 | 35.5 | 38.2 | 37.9 | |
| Ours-RL-7BModel Category=Our Models, Training Strategy=RLVR, Parameter Scale=7B2026.07 | 39.1 | 41.8 | 28.2 | 36.4 | |
| DiagAgent-14BModel Category=Medical Reasoning Models, Parameter Scale=14B2026.07 | 36.4 | 30 | 32.7 | 33 | |
| Qwen3-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 33.6 | 26.4 | 29.1 | 29.7 | |
| Qwen2.5-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 24.5 | 30.9 | 18.2 | 24.5 | |
| Huatuo-7BModel Category=Medical Reasoning Models, Parameter Scale=7B2026.07 | 23.6 | 18.2 | 17.3 | 19.7 |