Medical Question Answering on CMExam (test)
91.84AccuracyDeepseek-v3.2-685B
Evaluation Results
| Method | Links | |
|---|---|---|
| Deepseek-v3.2-685BModel Group=Awesome General Models2026.01 | 91.84 | |
| DEEPMED-14B-RLModel Group=DeepResearch Models, Training Stage=RL2026.01 | 91.19 | |
| DEEPMED-14B-SFTModel Group=DeepResearch Models, Training Stage=SFT2026.01 | 90.31 | |
| QuarkMed-32BModel Group=Medical Reasoning Models2026.01 | 88.61 | |
| Kimi-K2-Thinking-1TBModel Group=Awesome General Models2026.01 | 88.47 | |
| Tongyi-DeepResearch-30BA3BModel Group=DeepResearch Models2026.01 | 87.59 | |
| Gemini2.5-ProModel Group=Awesome General Models2026.01 | 87.4 | |
| Qwen3-30BA3B-ThinkingModel Group=Awesome General Models2026.01 | 83.81 | |
| Qwen3-14BModel Group=Awesome General Models2026.01 | 82.79 | |
| M1-1K-32BModel Group=Medical Reasoning Models2026.01 | 79.03 | |
| BaiChuan-M2-32BModel Group=Medical Reasoning Models2026.01 | 78.73 | |
| SafeMed-R1Optimization=CoT reasoning, safety-ethics alignment, improved GRPO2026.05 | 78.18 | |
| Qwen3-32B+SFT2026.05 | 77.11 | |
| qwen3-30B-A3B-Instruct-25072026.05 | 76.64 | |
| Qwen3-32B2026.05 | 76.38 | |
| Baichuan-M2-32B2026.05 | 75.89 | |
| gemini-2.5-flash2026.05 | 72.79 | |
| M1-1K-7BModel Group=Medical Reasoning Models2026.01 | 70.64 | |
| HuatuoGPT-o1-70BModel Group=Medical Reasoning Models2026.01 | 70.42 | |
| MedResaon-8BModel Group=Medical Reasoning Models2026.01 | 68.95 | |
| DeepSeek-R1-Distill-Qwen-32B2026.05 | 67.29 | |
| Qwen2.5-32B-Instruct2026.05 | 66.83 | |
| gpt-5-mini2026.05 | 66.17 | |
| Lingshu-32B2026.05 | 58.17 | |
| HuatuoGPT-o1-70B2026.05 | 55.37 | |
| google_medgemma-27b-text-it2026.05 | 54.56 |