Language Modeling on LM (CE 128-255 tokens)
2.69CE (128-255 tokens)Mamba-130m-hf
Evaluation Results
| Method | Links | |
|---|---|---|
| Mamba-130m-hfcontext_type=Recurrent models2026.03 | 2.69 | |
| GPT-2-124mcontext_type=Full context models (upper bound)2026.03 | 2.72 | |
| Pythia-160mcontext_type=Full context models (upper bound)2026.03 | 2.84 | |
| ARMT (GPT-2)context_type=Recurrent models2026.03 | 2.85 | |
| RMT (GPT-2)context_type=Recurrent models2026.03 | 2.91 | |
| GradMem (GPT-2, K = 1)K=1, Number of memory tokens=322026.03 | 2.92 | |
| GPT-2-124mLimit context to 128 tokens=true2026.03 | 3.2 |