Medical Question Answering on MedQA TW
88.89Accuracygpt-5-mini
Evaluation Results
| Method | Links | |
|---|---|---|
| gpt-5-mini2026.05 | 88.89 | |
| gemini-2.5-flash2026.05 | 88.61 | |
| SafeMed-R1Optimization=CoT reasoning, safety-ethics alignment, improved GRPO2026.05 | 85.14 | |
| Baichuan-M2-32B2026.05 | 84.43 | |
| Qwen3-32B+SFT2026.05 | 82.59 | |
| Qwen3-32B2026.05 | 81.53 | |
| qwen3-30B-A3B-Instruct-25072026.05 | 78.98 | |
| DeepSeek-R1-Distill-Qwen-32B2026.05 | 73.53 | |
| HuatuoGPT-o1-70B2026.05 | 73.32 | |
| Qwen2.5-32B-Instruct2026.05 | 67.16 | |
| google_medgemma-27b-text-it2026.05 | 65.62 | |
| Lingshu-32B2026.05 | 54.49 |