Instruction Following on Vicuna (Rouge-L)
20.93Rouge-LTeacher (Qwen1.5)
Evaluation Results
| Method | Links | |
|---|---|---|
| Teacher (Qwen1.5)Distillation Setting=Cross-Tokenizer KD2026.03 | 20.93 | |
| ARMADATeacher=text-to-audio, Distillation Type=L_cosine2026.03 | 19.96 | |
| ARMADATeacher=text-to-audio, Distillation Type=L_euclid2026.03 | 19.51 | |
| ARMADATeacher=text-to-video, Distillation Type=L_euclid2026.03 | 18.94 | |
| ARMADATeacher=text-to-image, Distillation Type=L_cosine2026.03 | 18.89 | |
| ARMADATeacher=text-to-video, Distillation Type=L_cosine2026.03 | 18.82 | |
| ARMADATeacher=text-to-image, Distillation Type=L_euclid2026.03 | 18.72 | |
| ARMADATeacher=text-to-video, Distillation Type=L_elementwise2026.03 | 18.36 | |
| ARMADATeacher=text-to-audio, Distillation Type=L_elementwise2026.03 | 18.18 | |
| ARMADATeacher=text-to-image, Distillation Type=L_elementwise2026.03 | 18.16 | |
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 18.14 | |
| SFTModel=LLaMA-3, Role=Teacher, #Param=1.23B2026.04 | 18.01 | |
| LLaMA-3.1-8BDistillation Type=undistilled2026.03 | 17.85 | |
| DISTILLLENS + MiniLLMPolicy=Hybrid2026.02 | 17.8 | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 17.69 | |
| AMiDTeacher Model=GPT-2 XL (1.5B), Student Model=GPT-2 Medium (0.3B), Student Parameters=0.3B2025.10 | 17.69 | |
| CISTStudent=OPT 1.3B, Teacher=OPT 6.7B2026.05 | 17.69 | |
| OPT 6.7BRole=Teacher2026.05 | 17.64 | |
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 17.62 | |
| ABStudent=OPT 1.3B, Teacher=OPT 6.7B2026.05 | 17.59 | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 17.58 | |
| RKLStudent=OPT 1.3B, Teacher=OPT 6.7B2026.05 | 17.54 | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 17.45 | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 17.43 | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 17.36 | |
| SRKLStudent=OPT 1.3B, Teacher=OPT 6.7B2026.05 | 17.34 | |
| AKLStudent=OPT 1.3B, Teacher=OPT 6.7B2026.05 | 17.26 | |
| JSStudent=OPT 1.3B, Teacher=OPT 6.7B2026.05 | 17.15 | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 17.14 | |
| DistiLLMPolicy=On2026.02 | 17.1 | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 17.09 | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 17.02 | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 17.02 | |
| Sym-KLStudent=OPT 1.3B, Teacher=OPT 6.7B2026.05 | 16.95 | |
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 16.94 | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16.91 | |
| MiniLLMPolicy=On2026.02 | 16.9 | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 16.85 | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 16.76 | |
| AMiDTeacher Model=GPT-2 XL (1.5B), Student Model=GPT-2 (0.1B), Student Parameters=0.1B2025.10 | 16.76 | |
| FKLStudent=OPT 1.3B, Teacher=OPT 6.7B2026.05 | 16.71 | |
| SFTDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16.64 | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16.62 | |
| AMiDTeacher Model=GPT-2 XL (1.5B), Student Model=GPT-2 Large (0.8B), Student Parameters=0.8B2025.10 | 16.62 | |
| SFTDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 16.51 | |
| KDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16.43 | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 16.42 | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16.39 | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16.35 | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 16.34 | |
| GPT-2 1.5BRole=Teacher2026.05 | 16.31 | |
| Teacher (GPT-2 XL)Distillation Setting=Same-Tokenizer KD2026.03 | 16.24 | |
| DSKD-CMADistillation Setting=Cross-Tokenizer KD, Divergence Function=SRKL2026.03 | 16.21 | |
| TeacherModel=GPT-2 XL (1.5B)2025.10 | 16.12 | |
| GPT-2 XL (Teacher)Model Role=Teacher, Parameters=1.5B2025.10 | 16.12 | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 16.1 | |
| CISTStudent=GPT-2 0.1B, Teacher=GPT-2 1.5B2026.05 | 16.1 | |
| DSKD-CMA-CTDistillation Setting=Cross-Tokenizer KD, Divergence Function=KL2026.03 | 16.08 | |
| ImitKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 16 | |
| AKLDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 15.98 | |
| DSKD-CMA-GADistillation Setting=Cross-Tokenizer KD, Divergence Function=SKL2026.03 | 15.95 | |
| AKLDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 15.85 | |
| DSKD-CMA-GADistillation Setting=Cross-Tokenizer KD, Divergence Function=KL2026.03 | 15.84 | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 15.82 | |
| RKLStudent=GPT-2 0.1B, Teacher=GPT-2 1.5B2026.05 | 15.82 | |
| ABStudent=GPT-2 0.1B, Teacher=GPT-2 1.5B2026.05 | 15.82 | |
| DISTILLLENSPolicy=Off2026.02 | 15.8 | |
| Sym-KLStudent=GPT-2 0.1B, Teacher=GPT-2 1.5B2026.05 | 15.72 | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 15.71 | |
| GKDPolicy=On2026.02 | 15.7 | |
| MaKDModel=LLaMA-3, Role=Student, #Param=0.69B2026.04 | 15.62 | |
| KDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 15.59 | |
| MinEDDistillation Setting=Cross-Tokenizer KD2026.03 | 15.59 | |
| ImitKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 15.56 | |
| DSKDDistillation Setting=Same-Tokenizer KD2026.03 | 15.54 | |
| JSStudent=GPT-2 0.1B, Teacher=GPT-2 1.5B2026.05 | 15.47 | |
| ImitKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 15.32 | |
| DSKD-CLPDistillation Setting=Cross-Tokenizer KD2026.03 | 15.28 | |
| DSKD-CMADistillation Setting=Cross-Tokenizer KD, Divergence Function=KL2026.03 | 15.2 | |
| SRKLStudent=GPT-2 0.1B, Teacher=GPT-2 1.5B2026.05 | 14.96 | |
| AKLDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 14.94 | |
| KDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 14.93 | |
| DSKD-CLADistillation Setting=Cross-Tokenizer KD2026.03 | 14.93 | |
| ULDDistillation Setting=Cross-Tokenizer KD2026.03 | 14.85 | |
| SFTDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 14.79 | |
| KD (FKL)Policy=Off2026.02 | 14.7 | |
| SFTDistillation Setting=Cross-Tokenizer KD2026.03 | 14.66 | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 14.6 | |
| FKLStudent=GPT-2 0.1B, Teacher=GPT-2 1.5B2026.05 | 14.53 | |
| SFTModel=GPT-2, Role=Teacher, #Param=124M2026.04 | 13.99 | |
| MaKDModel=GPT-2, Role=Student, #Param=77.4M2026.04 | 13.77 | |
| AKLStudent=GPT-2 0.1B, Teacher=GPT-2 1.5B2026.05 | 13.7 | |
| KDModel=LLaMA-3, Role=Student, #Param=0.69B2026.04 | 12.62 | |
| SeqKDModel=LLaMA-3, Role=Student, #Param=0.69B2026.04 | 12.54 | |
| SFTModel=GPT-2, Role=Student, #Param=77.4M2026.04 | 12.24 | |
| SeqKDModel=GPT-2, Role=Student, #Param=77.4M2026.04 | 12.14 | |
| KDModel=GPT-2, Role=Student, #Param=77.4M2026.04 | 11.82 | |
| SFTModel=LLaMA-3, Role=Student, #Param=0.69B2026.04 | 11.38 | |
| Student (GPT-2)Distillation Setting=Baseline2026.03 | 10.18 | |
| NoneModel=GPT-2, Role=Student, #Param=77.4M2026.04 | 9.51 |