Multiple Choice Question Answering on MMLU-Pro (Professional Medicine)
94AccuracyGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4oPrompting (shots)=0-shot2024.10 | 94 | |
| GPT-4oPrompting (shots)=5-shot2024.10 | 94 | |
| GPT-4T (May 2024)Prompting (shots)=0-shot2024.10 | 92 | |
| GPT-4T (May 2024)Prompting (shots)=5-shot2024.10 | 92 | |
| Gemini2.5-ProModel Group=Awesome General Models2026.01 | 84.3 | |
| Deepseek-v3.2-685BModel Group=Awesome General Models2026.01 | 83.19 | |
| Kimi-K2-Thinking-1TBModel Group=Awesome General Models2026.01 | 82.74 | |
| DEEPMED-14B-RLModel Group=DeepResearch Models, Training Stage=RL2026.01 | 79.93 | |
| AlphaMed-70BModel Group=Medical Reasoning Models2026.01 | 79.54 | |
| Tongyi-DeepResearch-30BA3BModel Group=DeepResearch Models2026.01 | 79.54 | |
| DEEPMED-14B-SFTModel Group=DeepResearch Models, Training Stage=SFT2026.01 | 77.92 | |
| Qwen3-30BA3B-ThinkingModel Group=Awesome General Models2026.01 | 76.94 | |
| HuatuoGPT-o1-70BModel Group=Medical Reasoning Models2026.01 | 76.61 | |
| BaiChuan-M2-32BModel Group=Medical Reasoning Models2026.01 | 75.05 | |
| Qwen3-14BModel Group=Awesome General Models2026.01 | 74.53 | |
| M1-1K-32BModel Group=Medical Reasoning Models2026.01 | 71.53 | |
| M1-1K-7BModel Group=Medical Reasoning Models2026.01 | 65.73 | |
| MedResaon-8BModel Group=Medical Reasoning Models2026.01 | 63.13 | |
| BioLinkBERT_largeNumber of Parameters=340M2022.03 | 50.7 | |
| UnifiedQANumber of Parameters=11B2022.03 | 43.2 | |
| GPT-3Number of Parameters=175B2022.03 | 38.7 |