Web-based Agent Interaction on WebShop
100CoT Match RateTeacher (OPT13B)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Teacher (OPT13B)Teacher Model=N/A, Backbone Architecture=OPT, Model Parameter Size=13B, Distillation Strategy=Teacher2025.05 | 100 | 73.2 | 35.9 | 5.9 | |
| Teacher (LLaMA13B)Teacher Model=N/A, Backbone Architecture=LLaMA, Model Parameter Size=13B, Distillation Strategy=Teacher2025.05 | 100 | 71.8 | 34.8 | 5.8 | |
| Teacher (Orca2-13B)Teacher Model=N/A, Backbone Architecture=Orca2, Model Parameter Size=13B, Distillation Strategy=Teacher2025.05 | 100 | 75.6 | 34.6 | 5.8 | |
| OPT-13BRole=Teacher2025.05 | 100 | — | — | — | |
| LLaMA-13BRole=Teacher2025.05 | 100 | — | — | — | |
| Structured Agent Distillation (Orca2-7B)Teacher Model=Orca2-13B, Backbone Architecture=Orca2, Model Parameter Size=7B, Distillation Strategy=Structured Agent Distillation2025.05 | 74.6 | 66.2 | 34.2 | 5.8 | |
| Structured Agent Distillation (LLaMA-7B)Teacher Model=LLaMA13B, Backbone Architecture=LLaMA, Model Parameter Size=7B, Distillation Strategy=Structured Agent Distillation2025.05 | 72.9 | 64.1 | 34.9 | 5.8 | |
| Structured Agent DistillationTeacher Model=LLaMA-13B, Student Model=LLaMA-7B2025.05 | 72.9 | — | — | — | |
| Structured Agent Distillation (OPT-6.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=6.7B, Distillation Strategy=Structured Agent Distillation2025.05 | 72.5 | 63.8 | 35.1 | 5.9 | |
| Structured Agent DistillationTeacher Model=OPT-13B, Student Model=OPT-6.7B2025.05 | 72.5 | — | — | — | |
| Token-Orca2-7BTeacher Model=Orca2-13B, Backbone Architecture=Orca2, Model Parameter Size=7B, Distillation Strategy=Token-level KD2025.05 | 70.2 | 61.7 | 36.9 | 6 | |
| SeqKD (Orca2-7B)Teacher Model=Orca2-13B, Backbone Architecture=Orca2, Model Parameter Size=7B, Distillation Strategy=SeqKD2025.05 | 69.4 | 60.4 | 37.2 | 6.1 | |
| KD (Orca2-7B)Teacher Model=Orca2-13B, Backbone Architecture=Orca2, Model Parameter Size=7B, Distillation Strategy=KD2025.05 | 68.7 | 59.2 | 37.8 | 6.2 | |
| Token-LLaMA-7BTeacher Model=LLaMA13B, Backbone Architecture=LLaMA, Model Parameter Size=7B, Distillation Strategy=Token-level KD2025.05 | 68.3 | 59.3 | 37.5 | 6.1 | |
| Token-levelTeacher Model=LLaMA-13B, Student Model=LLaMA-7B2025.05 | 68.3 | — | — | — | |
| Structured Agent Distillation (OPT-2.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=2.7B, Distillation Strategy=Structured Agent Distillation2025.05 | 67.9 | 56.4 | 36.2 | 6.1 | |
| Token-OPT-6.7BTeacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=6.7B, Distillation Strategy=Token-level KD2025.05 | 67.9 | 58.6 | 37.2 | 6.2 | |
| Structured Agent DistillationTeacher Model=OPT-13B, Student Model=OPT-2.7B2025.05 | 67.9 | — | — | — | |
| Token-levelTeacher Model=OPT-13B, Student Model=OPT-6.7B2025.05 | 67.9 | — | — | — | |
| SeqKD (LLaMA-7B)Teacher Model=LLaMA13B, Backbone Architecture=LLaMA, Model Parameter Size=7B, Distillation Strategy=SeqKD2025.05 | 67.1 | 58.1 | 37.9 | 6.2 | |
| SeqKD (OPT-6.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=6.7B, Distillation Strategy=SeqKD2025.05 | 67 | 56.9 | 37.8 | 6.3 | |
| KD (OPT-6.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=6.7B, Distillation Strategy=KD2025.05 | 66.5 | 55.8 | 38.4 | 6.4 | |
| KD (LLaMA-7B)Teacher Model=LLaMA13B, Backbone Architecture=LLaMA, Model Parameter Size=7B, Distillation Strategy=KD2025.05 | 66 | 56.9 | 38.3 | 6.3 | |
| SeqKDTeacher Model=LLaMA-13B, Student Model=LLaMA-7B2025.05 | 65 | — | — | — | |
| Structured Agent Distillation (OPT-1.3B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=1.3B, Distillation Strategy=Structured Agent Distillation2025.05 | 63.8 | 48.7 | 38 | 6.4 | |
| Structured Agent DistillationTeacher Model=OPT-13B, Student Model=OPT-1.3B2025.05 | 63.8 | — | — | — | |
| SeqKDTeacher Model=OPT-13B, Student Model=OPT-6.7B2025.05 | 63.8 | — | — | — | |
| KDTeacher Model=LLaMA-13B, Student Model=LLaMA-7B2025.05 | 63.5 | — | — | — | |
| Token-OPT-2.7BTeacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=2.7B, Distillation Strategy=Token-level KD2025.05 | 62.7 | 51 | 39 | 6.6 | |
| Token-levelTeacher Model=OPT-13B, Student Model=OPT-2.7B2025.05 | 62.7 | — | — | — | |
| SeqKD (OPT-2.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=2.7B, Distillation Strategy=SeqKD2025.05 | 61.9 | 49.7 | 39.5 | 6.7 | |
| KDTeacher Model=OPT-13B, Student Model=OPT-6.7B2025.05 | 61.5 | — | — | — | |
| KD (OPT-2.7B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=2.7B, Distillation Strategy=KD2025.05 | 61 | 48.3 | 40.1 | 6.9 | |
| SeqKDTeacher Model=OPT-13B, Student Model=OPT-2.7B2025.05 | 59.8 | — | — | — | |
| Token-OPT-1.3BTeacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=1.3B, Distillation Strategy=Token-level KD2025.05 | 57.9 | 43.2 | 41.8 | 7.1 | |
| Token-levelTeacher Model=OPT-13B, Student Model=OPT-1.3B2025.05 | 57.9 | — | — | — | |
| KDTeacher Model=OPT-13B, Student Model=OPT-2.7B2025.05 | 57.3 | — | — | — | |
| SeqKD (OPT-1.3B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=1.3B, Distillation Strategy=SeqKD2025.05 | 56.1 | 41.8 | 42.5 | 7.2 | |
| KD (OPT-1.3B)Teacher Model=OPT13B, Backbone Architecture=OPT, Model Parameter Size=1.3B, Distillation Strategy=KD2025.05 | 55 | 40.7 | 43.9 | 7.3 | |
| SeqKDTeacher Model=OPT-13B, Student Model=OPT-1.3B2025.05 | 54.1 | — | — | — | |
| KDTeacher Model=OPT-13B, Student Model=OPT-1.3B2025.05 | 52.4 | — | — | — |