Medical Question Answering on MedQA (test)
92AccuracyDrugClaw
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DrugClawMode=graph2026.05 | 92 | — | — | — | |
| DrugClawMode=graph, web-search fallback=disabled2026.05 | 91.7 | — | — | — | |
| DrugClawMode=linear2026.05 | 91 | — | — | — | |
| Biomni2026.05 | 90.2 | — | — | — | |
| DeepEvidence2026.05 | 90.2 | — | — | — | |
| Direct LLMModel=GPT-52026.05 | 89.8 | — | — | — | |
| Direct LLMModel=GPT-5-mini2026.05 | 88.4 | — | — | — | |
| ToolUniverse2026.05 | 88 | — | — | — | |
| RLSRτ=1.02026.07 | 78.9 | — | — | 3 | |
| RLCRτ=0.952026.07 | 72.2 | — | — | 1.4 | |
| OursBase Model=Qwen2.5-7B-Instruct, Fine-tuning Dataset Source=m23k, Selection Method=Ours, Number of Selected Examples=20002026.06 | 71.4 | — | — | — | |
| Middle PerplexityBase Model=Qwen2.5-7B-Instruct, Fine-tuning Dataset Source=m23k, Selection Method=Middle Perplexity, Number of Selected Examples=20002026.06 | 69.7 | — | — | — | |
| LearnabilityBase Model=Qwen2.5-7B-Instruct, Fine-tuning Dataset Source=m23k, Selection Method=Learnability, Number of Selected Examples=20002026.06 | 69 | — | — | — | |
| RandomBase Model=Qwen2.5-7B-Instruct, Fine-tuning Dataset Source=m23k, Selection Method=Random, Number of Selected Examples=20002026.06 | 68.8 | — | — | — | |
| Embedding DiversityBase Model=Qwen2.5-7B-Instruct, Fine-tuning Dataset Source=m23k, Selection Method=Embedding Diversity, Number of Selected Examples=20002026.06 | 68.8 | — | — | — | |
| S2LBase Model=Qwen2.5-7B-Instruct, Fine-tuning Dataset Source=m23k, Selection Method=S2L, Number of Selected Examples=20002026.06 | 67.6 | — | — | — | |
| IDCBackbone=Llama-3.1-8B, Method Type=Iterative Distractor Curation2026.03 | 66.14 | — | — | — | |
| DirectBackbone=Llama-3.1-8B, Method Type=Direct conversion to short-answer2026.03 | 65.2 | — | — | — | |
| FilterBackbone=Llama-3.1-8B, Method Type=Filtering non-convertible questions2026.03 | 65.04 | — | — | — | |
| RewriteBackbone=Llama-3.1-8B, Method Type=Model-based conversion2026.03 | 64.41 | — | — | — | |
| Orig.Backbone=Llama-3.1-8B, Method Type=Standard RLVR on original MCQs2026.03 | 62.14 | — | — | — | |
| BaseBackbone=Llama-3.1-8B, Method Type=Base instruct model2026.03 | 59.07 | — | — | — | |
| IDCBackbone=Qwen2-7B, Method Type=Iterative Distractor Curation2026.03 | 49.8 | — | — | — | |
| Orig.Backbone=Qwen2-7B, Method Type=Standard RLVR on original MCQs2026.03 | 49.25 | — | — | — | |
| DirectBackbone=Qwen2-7B, Method Type=Direct conversion to short-answer2026.03 | 48.7 | — | — | — | |
| RewriteBackbone=Qwen2-7B, Method Type=Model-based conversion2026.03 | 48.39 | — | — | — | |
| FilterBackbone=Qwen2-7B, Method Type=Filtering non-convertible questions2026.03 | 48 | — | — | — | |
| BaseBackbone=Qwen2-7B, Method Type=Base instruct model2026.03 | 47.21 | — | — | — | |
| Accuracy Rankingn=71, Domain (Format)=Med Knowledge (MCQ)2025.09 | — | — | 0.893 | — | |
| EAGLE2model=Qwen2.5-14B, Temperature=02025.03 | — | 2.46 | 2.57 | — | |
| EAGLE2model=LLaMA3-instruct-8B, Temperature=02025.03 | — | 3.13 | 2.76 | — | |
| EAGLE2model=Qwen2.5-14B, Temperature=12025.03 | — | 2.24 | 2.06 | — | |
| EAGLE2model=LLaMA3-instruct-8B, Temperature=12025.03 | — | 2.24 | 2.47 | — | |
| IRT (pr.) Rankingn=71, Domain (Format)=Med Knowledge (MCQ)2025.09 | — | — | 0.899 | — | |
| IRT(All) Rankingn=71, Domain (Format)=Med Knowledge (MCQ)2025.09 | — | — | 0.901 | — | |
| RASD(REST)model=Qwen2.5-14B, Temperature=02025.03 | — | 2.57 | 2.94 | — | |
| RASD(REST)model=LLaMA3-instruct-8B, Temperature=02025.03 | — | 3.26 | 3 | — | |
| RASD(REST)model=Qwen2.5-14B, Temperature=12025.03 | — | 2.35 | 2.36 | — | |
| RASD(REST)model=LLaMA3-instruct-8B, Temperature=12025.03 | — | 2.35 | 2.75 | — | |
| RESTmodel=Qwen2.5-14B, Temperature=02025.03 | — | 1.66 | 0.84 | — | |
| RESTmodel=LLaMA3-instruct-8B, Temperature=02025.03 | — | 1.85 | 0.89 | — | |
| RESTmodel=Qwen2.5-14B, Temperature=12025.03 | — | 1.42 | 0.5 | — | |
| RESTmodel=LLaMA3-instruct-8B, Temperature=12025.03 | — | 1.3 | 0.51 | — |