Speculative Decoding on Pre-training dataset
0.658Speculative Accept %Full Knowledge Distillation
Evaluation Results
| Method | Links | |
|---|---|---|
| Full Knowledge DistillationTraining Tokens=100B, Student Model=300M, Teacher Model=3B2025.03 | 0.658 | |
| Random Sampling KDTraining Tokens=100B, Student Model=300M, Teacher Model=3B2025.03 | 0.657 | |
| Cross-EntropyTraining Tokens=100B, Student Model=300M, Teacher Model=3B2025.03 | 0.646 |