Instruction Following on DollyEval
32.5Rouge-LOpenLLaMA2-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| OpenLLaMA2-7BBackbone=OpenLLaMA2-7B, Role=Teacher, Training=Fine-tuned2024.06 | 32.5 | 58.8 | — | |
| Adversarial Moment-Matching DistillationBackbone=OpenLLaMA-3B, Role=Student, Training=Fine-tuned2024.06 | 30.7 | 59.8 | — | |
| TeacherBackbone=LLaMA, Parameters=13B2023.06 | 29.7 | 79 | — | |
| DistiLLMBackbone=OpenLLaMA-3B, Role=Student, Training=Fine-tuned2024.06 | 29.5 | 59.2 | — | |
| TeacherBackbone=OPT, Parameters=13B2023.06 | 29.2 | 70.3 | — | |
| MINILLMBackbone=OPT, Parameters=6.7B2023.06 | 29 | 70.8 | — | |
| MINILLMBackbone=LLaMA, Parameters=7B2023.06 | 29 | 76.4 | — | |
| SeqKDBackbone=OPT, Parameters=6.7B2023.06 | 28.5 | 69.6 | — | |
| MiniLLMStudent Model=GPT-Neo-2.7B, Teacher Model=GPT-J-6B2023.06 | 28.5 | 63.4 | — | |
| MiniLLMBackbone=OpenLLaMA-3B, Role=Student, Training=Fine-tuned2024.06 | 28.4 | 58.7 | — | |
| KDBackbone=OPT, Parameters=6.7B2023.06 | 28.3 | 68.6 | — | |
| GPT-2 XL (teacher)Model Size=1.5B2024.06 | 28.2 | 45.5 | — | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 27.86 | — | — | |
| AMiDTeacher Model=GPT-2 XL (1.5B), Student Model=GPT-2 Large (0.8B), Student Parameters=0.8B2025.10 | 27.86 | — | — | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 27.67 | — | — | |
| TeacherBackbone=GPT-2, Parameters=1.5B2023.06 | 27.6 | 58.4 | — | |
| SFT w/o KDBackbone=OPT, Parameters=6.7B2023.06 | 27.6 | 67.9 | — | |
| SFT w/o KDStudent Model=GPT-2-1.5B, Teacher Model=GPT-J-6B2023.06 | 27.6 | 58.4 | — | |
| SeqKDBackbone=OPT, Parameters=2.7B2023.06 | 27.5 | 57.6 | — | |
| SeqKDBackbone=LLaMA, Parameters=7B2023.06 | 27.5 | 73.6 | — | |
| GKDBackbone=OpenLLaMA-3B, Role=Student, Training=Fine-tuned2024.06 | 27.5 | 57.6 | — | |
| MINILLMBackbone=OPT, Parameters=2.7B2023.06 | 27.4 | 63.2 | — | |
| KDBackbone=LLaMA, Parameters=7B2023.06 | 27.4 | 73.7 | — | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 27.34 | — | — | |
| AMiDTeacher Model=GPT-2 XL (1.5B), Student Model=GPT-2 Medium (0.3B), Student Parameters=0.3B2025.10 | 27.34 | — | — | |
| TeacherModel=GPT-J-6B2023.06 | 27.3 | 65.8 | — | |
| TeacherModel=GPT-2 XL (1.5B)2025.10 | 27.14 | — | — | |
| GPT-2 XL (Teacher)Model Role=Teacher, Parameters=1.5B2025.10 | 27.14 | — | — | |
| SFT w/o KDBackbone=OPT, Parameters=2.7B2023.06 | 27.1 | 55.4 | — | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 27.09 | — | — | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 27.01 | — | — | |
| SeqKDStudent Model=GPT-2-1.5B, Teacher Model=GPT-J-6B2023.06 | 27 | 58.5 | — | |
| TeacherModel=GPT-2-1.5B2025.09 | 27 | — | — | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 26.93 | — | — | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 26.87 | — | — | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 26.85 | — | — | |
| SFT w/o KDStudent Model=GPT-Neo-2.7B, Teacher Model=GPT-J-6B2023.06 | 26.8 | 60.7 | — | |
| MINILLMBackbone=OPT, Parameters=1.3B2023.06 | 26.7 | 60.7 | — | |
| KDStudent Model=GPT-2-760M, Teacher Model=GPT-J-6B2023.06 | 26.7 | 51.6 | — | |
| KDStudent Model=GPT-Neo-2.7B, Teacher Model=GPT-J-6B2023.06 | 26.7 | 61.5 | — | |
| SFTBackbone=OpenLLaMA-3B, Role=Student, Training=Fine-tuned2024.06 | 26.7 | 46.8 | — | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 26.61 | — | — | |
| KDStudent Model=GPT-2-1.5B, Teacher Model=GPT-J-6B2023.06 | 26.6 | 56.5 | — | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 26.5 | — | — | |
| AMiDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 26.44 | — | — | |
| AMiDTeacher Model=GPT-2 XL (1.5B), Student Model=GPT-2 (0.1B), Student Parameters=0.1B2025.10 | 26.44 | — | — | |
| MINILLMBackbone=GPT-2, Parameters=760M2023.06 | 26.4 | 54.7 | — | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 26.38 | — | — | |
| SFT w/o KDBackbone=LLaMA, Parameters=7B2023.06 | 26.3 | 73 | — | |
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 26.3 | — | — | |
| KDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 26.27 | — | — | |
| SeqKDBackbone=OpenLLaMA-3B, Role=Student, Training=Fine-tuned2024.06 | 26.2 | 50.2 | — | |
| SFTDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 26.17 | — | — | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 26.16 | — | — | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 26.12 | — | — | |
| SeqKDBackbone=OPT, Parameters=1.3B2023.06 | 26.1 | 51 | — | |
| MMKD (ours)Model Size=0.1B2024.06 | 26.1 | 31.7 | — | |
| SFT w/o KDBackbone=OPT, Parameters=1.3B2023.06 | 26 | 52.6 | — | |
| SeqKDStudent Model=GPT-2-760M, Teacher Model=GPT-J-6B2023.06 | 26 | 51.4 | — | |
| KDBackbone=GPT-2, Parameters=760M2023.06 | 25.9 | 53.4 | — | |
| KDBackbone=OPT, Parameters=2.7B2023.06 | 25.9 | 60.5 | — | |
| MiniLLMStudent Model=GPT-2-1.5B, Teacher Model=GPT-J-6B2023.06 | 25.9 | 59.6 | — | |
| MiniLLMStudent Model=GPT-2-760M, Teacher Model=GPT-J-6B2023.06 | 25.8 | 54 | — | |
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 25.8 | — | — | |
| TAIDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 25.74 | — | — | |
| DistiLLM (SRKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 25.74 | — | — | |
| SFTDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 25.7 | — | — | |
| SeqKDBackbone=GPT-2, Parameters=760M2023.06 | 25.6 | 52 | — | |
| SeqKDStudent Model=GPT-Neo-2.7B, Teacher Model=GPT-J-6B2023.06 | 25.6 | 60.8 | — | |
| AKLDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 25.57 | — | — | |
| SFT w/o KDBackbone=GPT-2, Parameters=340M2023.06 | 25.5 | 51.9 | — | |
| DistiLLM (SKL)Distillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 25.5 | — | — | |
| ABKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 25.49 | — | — | |
| AKLDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Large (0.8B), Student Model Parameters=0.8B2025.10 | 25.45 | — | — | |
| MINILLMBackbone=GPT-2, Parameters=340M2023.06 | 25.4 | 52.2 | — | |
| SFT w/o KDBackbone=GPT-2, Parameters=760M2023.06 | 25.4 | 50.7 | — | |
| KDBackbone=OPT, Parameters=1.3B2023.06 | 25.4 | 52.7 | — | |
| SFT w/o KDStudent Model=GPT-2-760M, Teacher Model=GPT-J-6B2023.06 | 25.4 | 50.7 | — | |
| SeqKDBackbone=GPT-2, Parameters=340M2023.06 | 25.3 | 50.5 | — | |
| ImitKDBackbone=OpenLLaMA-3B, Role=Student, Training=Fine-tuned2024.06 | 25.3 | 53.7 | — | |
| DistiLLMModel Size=0.1B2024.06 | 25.2 | 31.2 | — | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 25.06 | — | — | |
| KDBackbone=GPT-2, Parameters=340M2023.06 | 25 | 51.6 | — | |
| CSDLoss=CSD2025.09 | 24.94 | — | — | |
| MINILLMBackbone=GPT-2, Parameters=120M2023.06 | 24.6 | 44.7 | — | |
| GKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 24.58 | — | — | |
| SRKLLoss=SRKL, Parameter=0.12025.09 | 24.53 | — | — | |
| MiniLLMDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 24.47 | — | — | |
| ImitKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 24.46 | — | — | |
| MiniLLMModel Size=0.1B2024.06 | 24.3 | 30.2 | — | |
| KDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 Medium (0.3B), Student Model Parameters=0.3B2025.10 | 24.27 | — | — | |
| RKLLoss=RKL2025.09 | 24.26 | — | — | |
| SeqKDModel Size=0.1B2024.06 | 24.2 | 29.8 | — | |
| ABLoss=AB, Parameters=0.2, 0.72025.09 | 24.2 | — | — | |
| SeqKDDistillation Setup=GPT-2 XL (1.5B) → GPT-2 (0.1B), Student Model Parameters=0.1B2025.10 | 24.2 | — | — | |
| SKLLoss=SKL, Parameter=0.12025.09 | 24.17 | — | — | |
| GJSLoss=GJS, Parameter=0.92025.09 | 24.1 | — | — | |
| TVLoss=TV2025.09 | 23.88 | — | — | |
| KDModel Size=0.1B2024.06 | 23.8 | 29.5 | — | |
| GKDModel Size=0.1B2024.06 | 23.6 | 29.2 | — |