Instruction Following on Self-Instruct (ROUGE-L)
16.5ROUGE-LMiniLLM
Evaluation Results
| Method | Links | |
|---|---|---|
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16.5 | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16.46 | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 15.69 | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 15.46 | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 15.26 | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 15.07 | |
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 14.87 | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 14.61 | |
| TeacherModel=GPT-2 XL (1.5B)2025.10 | 14.55 | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 14.53 | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 14.44 | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 14.11 | |
| TeacherModel=GPT-2-1.5B2025.09 | 14.07 | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 13.93 | |
| AKLDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 13.83 | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 13.79 | |
| SFTDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 13.78 | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 13.74 | |
| KDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 13.72 | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 13.69 | |
| ImitKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 13.26 | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 13.01 | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 12.91 | |
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 12.83 | |
| SFTDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 12.6 | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 12.52 | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 12.36 | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 12.35 | |
| SRKLLoss=SRKL, Parameter=0.12025.09 | 12.19 | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 12.13 | |
| CSDLoss=CSD2025.09 | 12.06 | |
| AKLDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 12.06 | |
| ImitKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 12 | |
| ABLoss=AB, Parameters=0.2, 0.72025.09 | 11.82 | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 11.78 | |
| GJSLoss=GJS, Parameter=0.92025.09 | 11.4 | |
| SKLLoss=SKL, Parameter=0.12025.09 | 11.21 | |
| RKLLoss=RKL2025.09 | 11.19 | |
| AKLDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 11.18 | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 11.12 | |
| TVLoss=TV2025.09 | 11.03 | |
| JeffreyLoss=Jeffrey2025.09 | 10.82 | |
| KDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 10.58 | |
| ImitKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 10.34 | |
| Sym-KLLoss=Sym-KL2025.09 | 10.24 | |
| KDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 10.12 | |
| KLLoss=KL2025.09 | 10.02 | |
| SFTDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 9.62 |