Natural Language Understanding on NLU Benchmarks 0-shot
46.20-shot ScoreRandom Sampling KD
Evaluation Results
| Method | Links | |
|---|---|---|
| Random Sampling KDTraining Tokens=100B, Student Model=300M, Teacher Model=3B2025.03 | 46.2 | |
| Full Knowledge DistillationTraining Tokens=100B, Student Model=300M, Teacher Model=3B2025.03 | 46.2 | |
| Cross-EntropyTraining Tokens=100B, Student Model=300M, Teacher Model=3B2025.03 | 45 |