Language Modeling on 1.3B 26B-token pre-training corpus (val)
2.077Validation Cross-EntropyInterdomain Attention
Evaluation Results
| Method | Links | |
|---|---|---|
| Interdomain AttentionModel Size=1.3B, Training Tokens=26B2026.05 | 2.077 | |
| SoftmaxModel Size=1.3B, Training Tokens=26B2026.05 | 2.155 | |
| S4D-onlyModel Size=1.3B, Training Tokens=26B2026.05 | 2.212 |