Question Answering on MedQA (Accuracy)
71.72AccuracyAll
Evaluation Results
| Method | Links | |
|---|---|---|
| AllIteration=12026.04 | 71.72 | |
| EVOSELECTIteration=2, Selection Ratio=0.22026.04 | 71.64 | |
| AttributionIteration=1, Selection Ratio=0.52026.04 | 71.41 | |
| TSDSIteration=1, Selection Ratio=0.22026.04 | 71.35 | |
| RandomIteration=2, Selection Ratio=0.22026.04 | 71.33 | |
| EVOSELECTIteration=1, Selection Ratio=0.52026.04 | 71.09 | |
| EVOSELECTIteration=1, Selection Ratio=0.22026.04 | 71.01 | |
| Attr-DivIteration=2, Selection Ratio=0.22026.04 | 71.01 | |
| AllIteration=22026.04 | 70.93 | |
| AttributionIteration=1, Selection Ratio=0.22026.04 | 70.86 | |
| AttributionIteration=2, Selection Ratio=0.22026.04 | 70.86 | |
| EVOSELECTIteration=2, Selection Ratio=0.52026.04 | 70.86 | |
| Attr-DivIteration=2, Selection Ratio=0.52026.04 | 70.78 | |
| RandomIteration=1, Selection Ratio=0.22026.04 | 70.62 | |
| RandomIteration=1, Selection Ratio=0.52026.04 | 70.62 | |
| RandomIteration=2, Selection Ratio=0.52026.04 | 70.62 | |
| DiversityIteration=1, Selection Ratio=0.22026.04 | 70.54 | |
| Attr-DivIteration=1, Selection Ratio=0.22026.04 | 70.54 | |
| DiversityIteration=2, Selection Ratio=0.22026.04 | 70.54 | |
| Attr-DivIteration=1, Selection Ratio=0.52026.04 | 70.38 | |
| TSDSIteration=2, Selection Ratio=0.22026.04 | 70.38 | |
| AttributionIteration=2, Selection Ratio=0.52026.04 | 70.31 | |
| DiversityIteration=2, Selection Ratio=0.52026.04 | 70.31 | |
| Base2026.04 | 70.23 | |
| DiversityIteration=1, Selection Ratio=0.52026.04 | 69.99 | |
| TSDSIteration=1, Selection Ratio=0.52026.04 | 69.91 | |
| TSDSIteration=2, Selection Ratio=0.52026.04 | 69.68 | |
| POPBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 56.47 | |
| POPModel=Qwen-2.5-7B-Inst, Training Strategy=POP, Evaluation Protocol=0-shot2026.04 | 56.08 | |
| BaseBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 55.94 | |
| BaseModel=Qwen-2.5-7B-Inst, Training Strategy=Base, Evaluation Protocol=0-shot2026.04 | 55.94 | |
| Train on DBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 55.3 | |
| Train on DBackbone=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | 53.65 | |
| EVOSELECTIteration=1, Selection Ratio=0.5, Backbone=3b base model2026.04 | 53.34 | |
| Train on DModel=Qwen-2.5-7B-Inst, Training Strategy=Naive pretraining on D, Evaluation Protocol=0-shot2026.04 | 53.33 | |
| Attr-DivIteration=1, Selection Ratio=0.2, Backbone=3b base model2026.04 | 53.02 | |
| POPBackbone=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | 52.98 | |
| BaseBackbone=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | 52.95 | |
| BaseModel=Qwen-2.5-7B, Training Strategy=Base, Evaluation Protocol=0-shot2026.04 | 52.95 | |
| POPModel=Qwen-2.5-7B, Training Strategy=POP, Evaluation Protocol=0-shot2026.04 | 52.95 | |
| RandomIteration=1, Selection Ratio=0.2, Backbone=3b base model2026.04 | 52.63 | |
| AttributionIteration=1, Selection Ratio=0.2, Backbone=3b base model2026.04 | 52.32 | |
| DiversityIteration=2, Selection Ratio=0.2, Backbone=3b base model2026.04 | 52.24 | |
| RandomIteration=1, Selection Ratio=0.5, Backbone=3b base model2026.04 | 52.08 | |
| DiversityIteration=1, Selection Ratio=0.5, Backbone=3b base model2026.04 | 51.85 | |
| EVOSELECTIteration=1, Selection Ratio=0.2, Backbone=3b base model2026.04 | 51.69 | |
| EVOSELECTIteration=2, Selection Ratio=0.5, Backbone=3b base model2026.04 | 51.61 | |
| EVOSELECTIteration=2, Selection Ratio=0.2, Backbone=3b base model2026.04 | 51.53 | |
| TSDSIteration=2, Selection Ratio=0.2, Backbone=3b base model2026.04 | 51.45 | |
| TSDSIteration=1, Selection Ratio=0.2, Backbone=3b base model2026.04 | 51.37 | |
| Attr-DivIteration=1, Selection Ratio=0.5, Backbone=3b base model2026.04 | 51.06 | |
| AttributionIteration=2, Selection Ratio=0.5, Backbone=3b base model2026.04 | 50.82 | |
| AttributionIteration=1, Selection Ratio=0.5, Backbone=3b base model2026.04 | 50.67 | |
| TSDSIteration=2, Selection Ratio=0.5, Backbone=3b base model2026.04 | 50.67 | |
| BaseIteration=Base, Backbone=3b base model2026.04 | 50.43 | |
| TSDSIteration=1, Selection Ratio=0.5, Backbone=3b base model2026.04 | 50.43 | |
| DiversityIteration=1, Selection Ratio=0.2, Backbone=3b base model2026.04 | 50.35 | |
| AllIteration=2, Selection Ratio=All, Backbone=3b base model2026.04 | 50.35 | |
| RandomIteration=2, Selection Ratio=0.2, Backbone=3b base model2026.04 | 50.35 | |
| Attr-DivIteration=2, Selection Ratio=0.2, Backbone=3b base model2026.04 | 50.27 | |
| RandomIteration=2, Selection Ratio=0.5, Backbone=3b base model2026.04 | 50.2 | |
| AttributionIteration=2, Selection Ratio=0.2, Backbone=3b base model2026.04 | 50.2 | |
| DiversityIteration=2, Selection Ratio=0.5, Backbone=3b base model2026.04 | 49.65 | |
| Attr-DivIteration=2, Selection Ratio=0.5, Backbone=3b base model2026.04 | 49.33 | |
| AllIteration=1, Selection Ratio=All, Backbone=3b base model2026.04 | 49.18 | |
| Train on DModel=Qwen-2.5-7B, Training Strategy=Naive pretraining on D, Evaluation Protocol=0-shot2026.04 | 49.13 | |
| T5 (beam)QA Model=T52026.04 | 32.68 | |
| T5 (beam)QA Model=BART2026.04 | 31.81 | |
| GPT-3 (clustering)QA Model=T52026.04 | 29.69 | |
| GPT-3 (clustering)QA Model=BART2026.04 | 28.04 | |
| T5 (contrast)QA Model=T52026.04 | 26.71 | |
| TinyLlamaQA Model=BART2026.04 | 25.22 | |
| T5 (contrast)QA Model=BART2026.04 | 24.19 | |
| TinyLlamaQA Model=T52026.04 | 24.12 | |
| Ground-TruthQA Model=T52026.04 | 21.76 | |
| GPT-3 (k-NN)QA Model=T52026.04 | 21.05 | |
| Ground-TruthQA Model=BART2026.04 | 20.42 | |
| SFT2026.04 | 20.4 | |
| GPT-3 (k-NN)QA Model=BART2026.04 | 20.03 | |
| GPT-3 (COT-k-NN)QA Model=T52026.04 | 19.95 | |
| R-Tuning2026.04 | 19.7 | |
| GPT-3 (COT-k-NN)QA Model=BART2026.04 | 19.25 | |
| KWT (LLM)Knowledge Estimation=LLM2026.04 | 18.5 | |
| FT-TOP2026.04 | 17.2 | |
| KWT (EM)Knowledge Estimation=EM2026.04 | 16.9 | |
| KWT (Rouge)Knowledge Estimation=Rouge2026.04 | 16.9 |