Multiple Choice Question Answering on Pulmonary Choice Accuracy KGQA (test)
67.6Overall AccuracyLung-R1-14B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Lung-R1-14BSize=14B, Training=KG-guided RL2026.06 | 67.6 | 81.72 | 26.56 | |
| Gemini-3-ProVariant=Pro2026.06 | 67.2 | 77.96 | 35.94 | |
| Lung-R1-7BSize=7B, Training=KG-guided RL2026.06 | 67.2 | 77.96 | 35.94 | |
| DeepSeek-V3.2Version=3.22026.06 | 66.8 | 77.96 | 34.38 | |
| Qwen3-30B-A3BSize=30B, Architecture=A3B2026.06 | 64.8 | 74.19 | 37.5 | |
| DeepSeek-R1Variant=R12026.06 | 63.6 | 71.51 | 40.62 | |
| Qwen2.5-14BSize=14B2026.06 | 63.2 | 74.19 | 31.25 | |
| HuatuoGPT-o1-7BSize=7B2026.06 | 62.8 | 77.96 | 18.75 | |
| Baichuan-M2Version=M22026.06 | 62.4 | 71.51 | 35.94 | |
| Claude-Sonnet-4.5Version=4.52026.06 | 60.4 | 67.2 | 40.62 | |
| ClinicalGPT-R1Variant=R12026.06 | 59.2 | 73.12 | 18.75 | |
| Qwen2.5-7BSize=7B2026.06 | 58.4 | 70.97 | 21.88 | |
| GPT-5.1Version=5.12026.06 | 54 | 61.29 | 32.81 | |
| GPT-5-chatVariant=chat2026.06 | 52 | 61.29 | 25 | |
| GPT-OSS-120BSize=120B2026.06 | 52 | 62.37 | 21.88 | |
| GPT-4.1Version=4.12026.06 | 50.4 | 59.68 | 23.44 | |
| GPT-4oVariant=o2026.06 | 49.6 | 58.06 | 25 | |
| GPT-OSS-20BSize=20B2026.06 | 47.6 | 59.14 | 14.06 | |
| MedGemma-27BSize=27B2026.06 | 45.2 | 55.38 | 15.62 | |
| Kimi-K2-ThinkingVariant=Thinking2026.06 | 44.4 | 55.38 | 12.5 |