Medical Diagnosis Accuracy on In-house Chinese Dataset (Follow-up Consultation)
54.3R1 ScoreOurs-SFT-32B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Ours-SFT-32BModel Category=Our Models, Training Strategy=SFT, Parameter Scale=32B2026.07 | 54.3 | 53.6 | 46 | 51.3 | |
| M2-32BModel Category=Medical Reasoning Models, Parameter Scale=32B2026.07 | 53.6 | 46.4 | 47.5 | 49.2 | |
| QwQ-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 51.8 | 52.2 | 47.1 | 50.4 | |
| Qwen2.5-72BModel Category=Chat models, Parameter Scale=72B2026.07 | 46.7 | 50 | 39.5 | 45.4 | |
| Qwen3-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 45.7 | 42 | 41.7 | 43.1 | |
| Ours-RL-7BModel Category=Our Models, Training Strategy=RLVR, Parameter Scale=7B2026.07 | 39.5 | 44.9 | 32.6 | 39 | |
| DiagAgent-14BModel Category=Medical Reasoning Models, Parameter Scale=14B2026.07 | 38 | 39.5 | 37.7 | 38.4 | |
| Qwen2.5-32BModel Category=Chat models, Parameter Scale=32B2026.07 | 36.6 | 37.7 | 28.3 | 34.2 | |
| Huatuo-7BModel Category=Medical Reasoning Models, Parameter Scale=7B2026.07 | 23.9 | 20.3 | 19.2 | 21.1 |