Medical LLM Evaluation on LLMEval Med
73.5Reasoning ScoreGemini-2.5-Pro
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Gemini-2.5-ProType=Closed-source LLM2026.02 | 73.5 | 72.5 | 60.1 | — | |
| Qwen3-235B-InstructType=Open-source LLM2026.02 | 65.7 | 76.2 | 61.5 | — | |
| Baichuan-M2-32BType=Specialized LLM2026.02 | 64.2 | 63.8 | 64.2 | — | |
| o3Type=Closed-source LLM2026.02 | 63.8 | 66.9 | 65.8 | — | |
| Deepseek-R1Type=Open-source LLM2026.02 | 63.4 | 69.6 | 69.6 | — | |
| GPT-4.1Type=Closed-source LLM2026.02 | 60.4 | 57.3 | 58.8 | — | |
| Qwen3-30B-A3B-Instruct + More Query Rubricsbackbone=Qwen3-30B-A3B-Instruct, supervision=More Query Rubrics2026.02 | 59.8 | 79.8 | 57.8 | — | |
| Qwen3-32BType=Open-source LLM2026.02 | 59.1 | 62.3 | 59.3 | — | |
| Claude-3.7-SonnetType=Closed-source LLM2026.02 | 57.8 | 75.2 | 54 | — | |
| HuatuoGPT-o1-72BType=Specialized LLM2026.02 | 56.9 | 56.3 | 49.5 | — | |
| Qwen3-30B-A3B-InstructType=Our Method baseline2026.02 | 54.9 | 67 | 52.1 | — | |
| Qwen3-4B-Instruct + Doctor Rubricsbackbone=Qwen3-4B-Instruct, supervision=Doctor Rubrics2026.02 | 46.1 | 81.7 | 55.9 | — | |
| Qwen3-4B-Instruct + More Query Rubricsbackbone=Qwen3-4B-Instruct, supervision=More Query Rubrics2026.02 | 44.7 | 87.2 | 55.4 | — | |
| Qwen3-4B-Instruct + Principle Rubricsbackbone=Qwen3-4B-Instruct, supervision=Principle Rubrics2026.02 | 42.2 | 80.7 | 52.6 | — | |
| Qwen3-4B-InstructType=Our Method baseline2026.02 | 39.2 | 66.1 | 49.8 | — | |
| Qwen3-4B-Instruct + Draft Rubricsbackbone=Qwen3-4B-Instruct, supervision=Draft Rubrics2026.02 | 39.2 | 85.3 | 52.6 | — | |
| DPOBackbone=Qwen2.5-7B-Instruct, Training Method=DPO2026.03 | — | — | — | 35.8 | |
| DPOBackbone=Llama-3.2-3B-Instruct, Training Method=DPO2026.03 | — | — | — | 11.5 | |
| DPOBackbone=Qwen3-4B-Instruct-2507, Training Method=DPO2026.03 | — | — | — | 74.9 | |
| HeRLBackbone=Qwen2.5-7B-Instruct, Training Method=HeRL2026.03 | — | — | — | 65 | |
| HeRLBackbone=Llama-3.2-3B-Instruct, Training Method=HeRL2026.03 | — | — | — | 18.7 | |
| HeRLBackbone=Qwen3-4B-Instruct-2507, Training Method=HeRL2026.03 | — | — | — | 79.3 | |
| Llama-3.2-3B-InstructBackbone=Llama-3.2-3B-Instruct, Training Method=Initial2026.03 | — | — | — | 16.1 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct, Training Method=Initial2026.03 | — | — | — | 56 | |
| Qwen3-4B-Instruct-2507Backbone=Qwen3-4B-Instruct-2507, Training Method=Initial2026.03 | — | — | — | 74.5 | |
| RLVRBackbone=Qwen2.5-7B-Instruct, Training Method=RLVR2026.03 | — | — | — | 60.5 | |
| RLVRBackbone=Llama-3.2-3B-Instruct, Training Method=RLVR2026.03 | — | — | — | 18.5 | |
| RLVRBackbone=Qwen3-4B-Instruct-2507, Training Method=RLVR2026.03 | — | — | — | 78.1 | |
| SFTBackbone=Qwen2.5-7B-Instruct, Training Method=SFT2026.03 | — | — | — | 34.8 | |
| SFTBackbone=Llama-3.2-3B-Instruct, Training Method=SFT2026.03 | — | — | — | 15.1 | |
| SFTBackbone=Qwen3-4B-Instruct-2507, Training Method=SFT2026.03 | — | — | — | 73.3 |