In-context recall on Recall-intensive tasks suite
62.3FDA Recall ScoreXfmr++
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Xfmr++Model parameters=2.7B, Training tokens=100B, L=32, d=2560, State size=N/A, Input truncation=2K tokens2024.09 | 62.3 | 30.9 | 44.3 | 29.3 | 61.8 | 21.4 | 41.7 | |
| Xfmr++Model parameters=1.3B, Training tokens=100B, L=24, d=2048, State size=N/A, Input truncation=2K tokens2024.09 | 46 | 29.2 | 41 | 24.8 | 58.8 | 21.3 | 36.9 | |
| GSAModel parameters=2.7B, Training tokens=100B, L=32, d=2560, State size=128 × Ld, Input truncation=2K tokens2024.09 | 39.1 | 33.5 | 39 | 26.9 | 60.8 | 19.9 | 36.5 | |
| GLAModel parameters=2.7B, Training tokens=100B, L=32, d=2560, State size=256 × Ld, Input truncation=2K tokens2024.09 | 30.3 | 35.5 | 36.8 | 23.3 | 58.2 | 21.8 | 34.3 | |
| GLAModel parameters=1.3B, Training tokens=100B, L=24, d=2048, State size=256 × Ld, Input truncation=2K tokens2024.09 | 26.7 | 30.6 | 34.8 | 21.5 | 56 | 19.1 | 31.4 | |
| RetNetModel parameters=2.7B, Training tokens=100B, L=32, d=2560, State size=512 × Ld, Input truncation=2K tokens2024.09 | 24.1 | 26.1 | 36.4 | 20.4 | 57.3 | 21.8 | 31 | |
| GSAModel parameters=1.3B, Training tokens=100B, L=24, d=2048, State size=128 × Ld, Input truncation=2K tokens2024.09 | 23.6 | 29.8 | 36 | 23.2 | 57 | 20.9 | 31.8 | |
| MambaModel parameters=2.7B, Training tokens=100B, L=32, d=2560, State size=64 × Ld, Input truncation=2K tokens2024.09 | 21.5 | 26.7 | 34.2 | 21.2 | 57 | 22.2 | 30.5 | |
| RetNetModel parameters=1.3B, Training tokens=100B, L=24, d=2048, State size=512 × Ld, Input truncation=2K tokens2024.09 | 21.2 | 27.2 | 34 | 15.5 | 52.7 | 20 | 28.4 | |
| HGRN2Model parameters=2.7B, Training tokens=100B, L=32, d=2560, State size=128 × Ld, Input truncation=2K tokens2024.09 | 15 | 29.9 | 35.1 | 17 | 59.8 | 20 | 29.5 | |
| MambaModel parameters=1.3B, Training tokens=100B, L=24, d=2048, State size=64 × Ld, Input truncation=2K tokens2024.09 | 13.9 | 25.4 | 33.2 | 18.5 | 53.5 | 21.7 | 27.7 | |
| HGRN2Model parameters=1.3B, Training tokens=100B, L=24, d=2048, State size=128 × Ld, Input truncation=2K tokens2024.09 | 9.9 | 23.1 | 32 | 16.4 | 55.2 | 19.1 | 25.9 |