Instruction Following on Evol-Inst
92.2Win RateTeacher
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TeacherModel=Qwen2.5-14B-Instruct2025.10 | 92.2 | — | |
| TeacherModel=Qwen2.5-7B-Instruct, Role=Teacher2025.10 | 89.6 | — | |
| DistiLLM-2α=-5, Teacher Model=Qwen2.5-14B-Instruct, Student Model=Qwen2.5-1.5B-Instruct2025.10 | 71.1 | — | |
| DistiLLM-2alpha=-5, Base Model=Qwen2.5-1.5B-Instruct2025.10 | 71 | — | |
| DistiLLM-2α=-1, Teacher Model=Qwen2.5-14B-Instruct, Student Model=Qwen2.5-1.5B-Instruct2025.10 | 70.4 | — | |
| MT (Teacher)Distillation Pair (Teacher Model -> Student Model)=Qwen2.5-7B-Inst -> Qwen2.5-1.5B2026.04 | 69.67 | — | |
| DistiLLM-2alpha=-1, Base Model=Qwen2.5-1.5B-Instruct2025.10 | 69 | — | |
| MT (Teacher)Distillation Pair (Teacher Model -> Student Model)=Qwen2-7B-Inst -> Qwen2-1.5B2026.04 | 59.75 | — | |
| VCRDDistillation Pair (Teacher Model -> Student Model)=Qwen2.5-7B-Inst -> Qwen2.5-1.5B2026.04 | 57.8 | — | |
| DistillLM-2Distillation Pair (Teacher Model -> Student Model)=Qwen2.5-7B-Inst -> Qwen2.5-1.5B2026.04 | 56.77 | — | |
| DistilLLMDistillation Pair (Teacher Model -> Student Model)=Qwen2.5-7B-Inst -> Qwen2.5-1.5B2026.04 | 55.38 | — | |
| KDDistillation Pair (Teacher Model -> Student Model)=Qwen2.5-7B-Inst -> Qwen2.5-1.5B2026.04 | 54.59 | — | |
| StudentModel=Qwen2.5-1.5B-Instruct2025.10 | 47.2 | — | |
| SeqKDDistillation Pair (Teacher Model -> Student Model)=Qwen2.5-7B-Inst -> Qwen2.5-1.5B2026.04 | 46.33 | — | |
| StudentModel=Qwen2.5-1.5B-Instruct, Role=Student2025.10 | 46.2 | — | |
| VCRDDistillation Pair (Teacher Model -> Student Model)=Qwen2-7B-Inst -> Qwen2-1.5B2026.04 | 44.49 | — | |
| DistillLM-2Distillation Pair (Teacher Model -> Student Model)=Qwen2-7B-Inst -> Qwen2-1.5B2026.04 | 37.27 | — | |
| DistilLLMDistillation Pair (Teacher Model -> Student Model)=Qwen2-7B-Inst -> Qwen2-1.5B2026.04 | 35.66 | — | |
| KDDistillation Pair (Teacher Model -> Student Model)=Qwen2-7B-Inst -> Qwen2-1.5B2026.04 | 32.45 | — | |
| MS (Student)Distillation Pair (Teacher Model -> Student Model)=Qwen2.5-7B-Inst -> Qwen2.5-1.5B2026.04 | 30.73 | — | |
| SeqKDDistillation Pair (Teacher Model -> Student Model)=Qwen2-7B-Inst -> Qwen2-1.5B2026.04 | 28.32 | — | |
| DPOOptimization Method=DPO, % Train=100%, Train set=Full, Backbone=LLaMA-3-8B2025.05 | 26.8 | 5.93 | |
| SimPOOptimization Method=SimPO, % Train=33%, Train set=Random, Backbone=LLaMA-3-8B2025.05 | 26.6 | 5.97 | |
| SimPOOptimization Method=SimPO, % Train=33%, Train set=LowAvg., Backbone=LLaMA-3-8B, Data Map Partition=Alignment Data Map2025.05 | 26.1 | 5.92 | |
| SimPOOptimization Method=SimPO, % Train=100%, Train set=Full, Backbone=LLaMA-3-8B2025.05 | 26.1 | 5.88 | |
| DPOOptimization Method=DPO, % Train=33%, Train set=LowAvg., Backbone=LLaMA-3-8B, Data Map Partition=Alignment Data Map2025.05 | 25.9 | 5.91 | |
| SimPOOptimization Method=SimPO, % Train=33%, Train set=HighVar., Backbone=LLaMA-3-8B, Data Map Partition=Alignment Data Map2025.05 | 25.9 | 5.94 | |
| DPOOptimization Method=DPO, % Train=33%, Train set=HighAvg., Backbone=LLaMA-3-8B, Data Map Partition=Alignment Data Map2025.05 | 25.7 | 5.97 | |
| SimPOOptimization Method=SimPO, % Train=33%, Train set=HighAvg., Backbone=LLaMA-3-8B, Data Map Partition=Alignment Data Map2025.05 | 25.5 | 5.94 | |
| DPOOptimization Method=DPO, % Train=33%, Train set=HighVar., Backbone=LLaMA-3-8B, Data Map Partition=Alignment Data Map2025.05 | 25.2 | 5.92 | |
| DPOOptimization Method=DPO, % Train=33%, Train set=Random, Backbone=LLaMA-3-8B2025.05 | 25 | 5.9 | |
| DPOOptimization Method=DPO, % Train=0%, Train set=Zeroshot, Backbone=LLaMA-3-8B2025.05 | 23.4 | 5.8 | |
| SimPOOptimization Method=SimPO, % Train=0%, Train set=Zeroshot, Backbone=LLaMA-3-8B2025.05 | 23.4 | 5.8 | |
| MS (Student)Distillation Pair (Teacher Model -> Student Model)=Qwen2-7B-Inst -> Qwen2-1.5B2026.04 | 19.49 | — |