Word Prediction on LAMBADA first 1000 examples (standard)
20.3Accuracy (LAMBADA)k=0 (baseline)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| k=0 (baseline)Model size=300M, Tokens=5B, Precision=FP16, Decoding=greedy, Mode=zero-shot, Examples=10002026.04 | 20.3 | 65.8 | |
| k=64 (DR only)Model size=300M, Tokens=5B, Precision=FP16, Decoding=greedy, Mode=zero-shot, Examples=1000, Depth Registers=k=642026.04 | 20.3 | 67.8 | |
| k=64 + sinkModel size=300M, Tokens=5B, Precision=FP16, Decoding=greedy, Mode=zero-shot, Examples=1000, Depth Registers=k=64, Attention-sink loss=included2026.04 | 20.3 | 64.7 |