Multi-hop Question Answering on HotPotQA (Task Success, CoT, Latency)
100CoT Match RateTeacher (OPT13B)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Teacher (OPT13B)Teacher Model=N/A, Backbone Architecture=OPT, Model Parameter Size=13B, Distillation Strategy=Teacher2025.05 | 100 | 82.7 | 40.7 | 4.8 | |
| Teacher (LLaMA13B)Teacher Model=N/A, Backbone Architecture=LLaMA, Model Parameter Size=13B, Distillation Strategy=Teacher2025.05 | 100 | 81 | 39.9 | 4.7 | |
| Teacher (Orca2-13B)Teacher Model=N/A, Backbone Architecture=Orca2, Model Parameter Size=13B, Distillation Strategy=Teacher2025.05 | 100 | 84.3 | 39.1 | 4.6 | |
| OPT-13BRole=Teacher2025.05 | 100 | — | — | — | |
| LLaMA-13BRole=Teacher2025.05 | 100 | — | — | — | |
| Teacher (GPT-2-1.5B)Model Scale=1.5B2025.05 | 100 | 78.5 | 10.8 | 6.2 | |
| Structured Agent Distillation (Orca2-7B)Teacher Model=Orca2-13B, Backbone Architecture=Orca2, Model Parameter Size=7B, Distillation Strategy=Structured Agent Distillation2025.05 | 86.5 | 78.6 | 38.9 | 4.7 | |
| Structured Agent Distillation (LLaMA-7B)Teacher Model=LLaMA13B, Backbone Architecture=LLaMA, Model Parameter Size=7B, Distillation Strategy=Structured Agent Distillation2025.05 | 84.7 | 75.2 | 39.8 | 4.8 | |
| Structured Agent DistillationTeacher Model=LLaMA-13B, Student Model=LLaMA-7B2025.05 | 84.7 | — | — | — | |
| Structured Agent Distillation (OPT-6.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=6.7B, Distillation Strategy=Structured Agent Distillation2025.05 | 84 | 73.9 | 40.2 | 4.9 | |
| Structured Agent DistillationTeacher Model=OPT-13B, Student Model=OPT-6.7B2025.05 | 84 | — | — | — | |
| Token-Orca2-7BTeacher Model=Orca2-13B, Backbone Architecture=Orca2, Model Parameter Size=7B, Distillation Strategy=Token-level KD2025.05 | 83.4 | 74.8 | 42 | 5.1 | |
| SeqKD (Orca2-7B)Teacher Model=Orca2-13B, Backbone Architecture=Orca2, Model Parameter Size=7B, Distillation Strategy=SeqKD2025.05 | 82.6 | 73.5 | 42.5 | 5.1 | |
| KD (Orca2-7B)Teacher Model=Orca2-13B, Backbone Architecture=Orca2, Model Parameter Size=7B, Distillation Strategy=KD2025.05 | 81.5 | 72.4 | 43.1 | 5.2 | |
| Token-LLaMA-7BTeacher Model=LLaMA13B, Backbone Architecture=LLaMA, Model Parameter Size=7B, Distillation Strategy=Token-level KD2025.05 | 81.2 | 71.5 | 43.2 | 5.2 | |
| Token-levelTeacher Model=LLaMA-13B, Student Model=LLaMA-7B2025.05 | 81.2 | — | — | — | |
| Structured Agent DistillationModel Scale=760M2025.05 | 80.4 | 73.1 | 11.7 | 6.6 | |
| Token-OPT-6.7BTeacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=6.7B, Distillation Strategy=Token-level KD2025.05 | 80.2 | 69.7 | 42.9 | 5.3 | |
| Token-levelTeacher Model=OPT-13B, Student Model=OPT-6.7B2025.05 | 80.2 | — | — | — | |
| SeqKD (LLaMA-7B)Teacher Model=LLaMA13B, Backbone Architecture=LLaMA, Model Parameter Size=7B, Distillation Strategy=SeqKD2025.05 | 80 | 70.9 | 43.6 | 5.2 | |
| Structured Agent Distillation (OPT-2.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=2.7B, Distillation Strategy=Structured Agent Distillation2025.05 | 79.8 | 67 | 41.7 | 5 | |
| Structured Agent DistillationTeacher Model=OPT-13B, Student Model=OPT-2.7B2025.05 | 79.8 | — | — | — | |
| SeqKD (OPT-6.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=6.7B, Distillation Strategy=SeqKD2025.05 | 79 | 68.1 | 43.3 | 5.4 | |
| KD (LLaMA-7B)Teacher Model=LLaMA13B, Backbone Architecture=LLaMA, Model Parameter Size=7B, Distillation Strategy=KD2025.05 | 79 | 70.1 | 44 | 5.3 | |
| KD (OPT-6.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=6.7B, Distillation Strategy=KD2025.05 | 78.1 | 67.2 | 43.9 | 5.4 | |
| SeqKDTeacher Model=LLaMA-13B, Student Model=LLaMA-7B2025.05 | 78 | — | — | — | |
| SeqKDTeacher Model=OPT-13B, Student Model=OPT-6.7B2025.05 | 77.1 | — | — | — | |
| KDTeacher Model=LLaMA-13B, Student Model=LLaMA-7B2025.05 | 76.5 | — | — | — | |
| Token-level KDModel Scale=760M2025.05 | 76.4 | 69.1 | 12 | 7.2 | |
| Token-OPT-2.7BTeacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=2.7B, Distillation Strategy=Token-level KD2025.05 | 75.5 | 62.9 | 45.2 | 5.6 | |
| Token-levelTeacher Model=OPT-13B, Student Model=OPT-2.7B2025.05 | 75.5 | — | — | — | |
| KDTeacher Model=OPT-13B, Student Model=OPT-6.7B2025.05 | 75.3 | — | — | — | |
| Structured Agent Distillation (OPT-1.3B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=1.3B, Distillation Strategy=Structured Agent Distillation2025.05 | 74.4 | 58.5 | 43.6 | 5.3 | |
| Structured Agent DistillationTeacher Model=OPT-13B, Student Model=OPT-1.3B2025.05 | 74.4 | — | — | — | |
| SeqKDModel Scale=760M2025.05 | 74.3 | 67.4 | 12.3 | 7.3 | |
| SeqKD (OPT-2.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=2.7B, Distillation Strategy=SeqKD2025.05 | 74.1 | 61.3 | 46 | 5.6 | |
| Structured Agent DistillationModel Scale=340M2025.05 | 74 | 65.5 | 12.2 | 7 | |
| KD (OPT-2.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=2.7B, Distillation Strategy=KD2025.05 | 73.4 | 60.7 | 46.5 | 5.7 | |
| KDModel Scale=760M2025.05 | 73.1 | 66.2 | 12.6 | 7.4 | |
| SeqKDTeacher Model=OPT-13B, Student Model=OPT-2.7B2025.05 | 72.3 | — | — | — | |
| Token-level KDModel Scale=340M2025.05 | 70.5 | 61.5 | 13 | 7.9 | |
| KDTeacher Model=OPT-13B, Student Model=OPT-2.7B2025.05 | 70.2 | — | — | — | |
| Token-OPT-1.3BTeacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=1.3B, Distillation Strategy=Token-level KD2025.05 | 69.8 | 54.1 | 48.5 | 6 | |
| Token-levelTeacher Model=OPT-13B, Student Model=OPT-1.3B2025.05 | 69.8 | — | — | — | |
| SeqKDModel Scale=340M2025.05 | 69.3 | 60.2 | 13.4 | 8.1 | |
| SeqKD (OPT-1.3B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=1.3B, Distillation Strategy=SeqKD2025.05 | 68.4 | 52.3 | 49.4 | 6.2 | |
| KDModel Scale=340M2025.05 | 68.3 | 58.8 | 13.8 | 8.2 | |
| KD (OPT-1.3B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=1.3B, Distillation Strategy=KD2025.05 | 67.3 | 51 | 50.1 | 6.3 | |
| Structured Agent DistillationModel Scale=120M2025.05 | 66.2 | 52.8 | 13.8 | 7.8 | |
| Token-level KDModel Scale=120M2025.05 | 65.7 | 48.3 | 14.8 | 8.9 | |
| SeqKDTeacher Model=OPT-13B, Student Model=OPT-1.3B2025.05 | 65.4 | — | — | — | |
| SeqKDModel Scale=120M2025.05 | 64.2 | 47.2 | 15 | 9 | |
| KDTeacher Model=OPT-13B, Student Model=OPT-1.3B2025.05 | 63.7 | — | — | — | |
| KDModel Scale=120M2025.05 | 63.4 | 46.1 | 15.3 | 9.2 |