Medical Reasoning on MMLU-Pro Health (English)
59.3AccuracyHuatuoGPT-o1-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| HuatuoGPT-o1-8BZero-shot=true2026.05 | 59.3 | |
| GPT-4oZero-shot=true2026.05 | 57.9 | |
| HiMed-8BBackbone=LLaMA-3.1-8B, Zero-shot=true2026.05 | 57.8 |
| Method | Links | |
|---|---|---|
| HuatuoGPT-o1-8BZero-shot=true2026.05 | 59.3 | |
| GPT-4oZero-shot=true2026.05 | 57.9 | |
| HiMed-8BBackbone=LLaMA-3.1-8B, Zero-shot=true2026.05 | 57.8 |