Long-context Language Modeling on LongPPL 32k
4.14Book PerplexityEngram-27B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Engram-27BPre-training Steps=50k, Pre-training Loss=1.62, Context Window Size=32k2026.01 | 4.14 | 2.82 | 2.44 | 13.41 | |
| Engram-27BPre-training Steps=46k, Pre-training Loss=1.63, Context Window Size=32k2026.01 | 4.19 | 2.84 | 2.45 | 13.59 | |
| Engram-27BPre-training Steps=41k, Pre-training Loss=1.66, Context Window Size=32k2026.01 | 4.37 | 2.92 | 2.5 | 14.26 | |
| MoE-27BPre-training Steps=50k, Pre-training Loss=1.63, Context Window Size=32k2026.01 | 4.38 | 2.91 | 2.49 | 14.16 |