Instruction Tuning Evaluation on ARC, GSM8k, HellaSwag, and MMLU (test val)
52.31ARC AccuracyLCG-DistilBERT-1k
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| LCG-DistilBERT-1kBase Model=Mistral-7B, Data Selection Strategy=Low-Confidence Gold (LCG), Classifier used for Selection=DistilBERT, Training Subset Size=1k2025.02 | 52.31 | 40 | 62.39 | 60.01 | 53.68 | |
| WizardLM-K-meansBase Model=Mistral-7B, Data Selection Strategy=K-means clustering2025.02 | 51.96 | 39.95 | 62.11 | 59.38 | 53.35 | |
| LCG-MNB-6kBase Model=Mistral-7B, Data Selection Strategy=Low-Confidence Gold (LCG), Classifier used for Selection=Multinomial Naive Bayes (MNB), Training Subset Size=6k2025.02 | 51.67 | 38.97 | 62.28 | 60.75 | 53.42 | |
| WizardLM-50k (Full)Base Model=Mistral-7B, Data Selection Strategy=Full dataset, Training Subset Size=50k2025.02 | 51.54 | 38.59 | 62.14 | 59.36 | 52.91 | |
| WizardLM-SuperFilteringBase Model=Mistral-7B, Data Selection Strategy=SuperFiltering2025.02 | 51.54 | 39.65 | 62.21 | 59.22 | 53.16 | |
| WizardLM-PerplexityBase Model=Mistral-7B, Data Selection Strategy=Perplexity-based filtering2025.02 | 51.19 | 39.77 | 62.21 | 59.01 | 53.05 | |
| WizardLM-LongestBase Model=Mistral-7B, Data Selection Strategy=Longest sequences2025.02 | 51.08 | 38.99 | 62.1 | 58.99 | 52.79 |