Language Modeling on WikiText-103
4.59PPLESPACE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ESPACEBackbone=Llama2-13B, Compression Ratio=20%2024.10 | 4.59 | — | |
| BaselineBackbone=Llama2-13B, Compression Ratio=0%2024.10 | 4.61 | — | |
| BaselineBackbone=Llama2-7B, Compression Ratio=0%2024.10 | 5.06 | — | |
| RetrainedBackbone=Llama2-7B, Compression Ratio=0%2024.10 | 5.06 | — | |
| ESPACEBackbone=Llama2-7B, Compression Ratio=21%2024.10 | 5.07 | — | |
| ESPACEBackbone=Llama2-13B, Compression Ratio=50%2024.10 | 5.13 | — | |
| ESPACEBackbone=Llama2-7B, Compression Ratio=50%2024.10 | 5.67 | — | |
| BaselineBackbone=Nemotron4-15B, Compression Ratio=0%2024.10 | 6.06 | — | |
| ESPACEBackbone=Nemotron4-15B, Compression Ratio=25%2024.10 | 6.28 | — | |
| ESPACEBackbone=GPT3-22B, Compression Ratio=40%2024.10 | 6.29 | — | |
| BaselineBackbone=GPT3-22B, Compression Ratio=0%2024.10 | 6.55 | — | |
| ESPACEBackbone=GPT3-22B, Compression Ratio=55%2024.10 | 6.73 | — | |
| ESPACEBackbone=Nemotron4-15B, Compression Ratio=50%2024.10 | 6.93 | — | |
| ESPACEBackbone=GPT3-8B, Compression Ratio=21%2024.10 | 7 | — | |
| BaselineBackbone=GPT3-8B, Compression Ratio=0%2024.10 | 7.38 | — | |
| ESPACEBackbone=GPT3-8B, Compression Ratio=50%2024.10 | 7.66 | — | |
| ADEPTArchitecture=Llama 7B, Reduction=10%2026.01 | 8.1 | — | |
| ADEPTArchitecture=Llama 7B, Reduction=20%2026.01 | 8.4 | — | |
| Base ModelArchitecture=Llama 7B, Reduction=None2026.01 | 8.6 | — | |
| MoA#total params=262M, Nheads=8, MACs=2.9G, Mem (floats)=9.9M2023.12 | 9.5 | — | |
| ESPACEBackbone=GPT3-1.3B, Compression Ratio=20%2024.10 | 9.53 | — | |
| SwitchHead#total params=262M, Nheads=2, MACs=2.0G, Mem (floats)=2.9M2023.12 | 9.55 | — | |
| Transformer#total params=262M, Nheads=16, MACs=5.4G, Mem (floats)=21.0M2023.12 | 9.66 | — | |
| MoA#total params=262M, Nheads=12, MACs=4.1G, Mem (floats)=14.7M2023.12 | 9.68 | — | |
| MoA#total params=262M, Nheads=4, MACs=1.7G, Mem (floats)=5.1M2023.12 | 9.69 | — | |
| MoA#total params=262M, Nheads=2, MACs=1.1G, Mem (floats)=2.7M2023.12 | 9.87 | — | |
| BaselineBackbone=GPT3-1.3B, Compression Ratio=0%2024.10 | 9.94 | — | |
| DeeBERTArchitecture=Llama 7B, Reduction=10%2026.01 | 10.1 | — | |
| PABEEArchitecture=Llama 7B, Reduction=10%2026.01 | 10.2 | — | |
| Hybrid H3Number of Parameters=2.7B2022.12 | 10.6 | — | |
| ESPACEBackbone=GPT3-1.3B, Compression Ratio=47%2024.10 | 11.07 | — | |
| GPT-NeoNumber of Parameters=2.7B2022.12 | 11.5 | — | |
| DeeBERTArchitecture=Llama 7B, Reduction=20%2026.01 | 12 | — | |
| PABEEArchitecture=Llama 7B, Reduction=20%2026.01 | 12.1 | — | |
| SwitchHead#total params=47M, Nheads=2, MACs=170.4M, Mem (floats)=0.8M2023.12 | 12.27 | — | |
| Transformer#total params=47M, Nheads=10, MACs=453.4M, Mem (floats)=3.5M2023.12 | 12.31 | — | |
| Hybrid H3Number of Parameters=1.3B2022.12 | 12.5 | — | |
| MoA#total params=47M, Nheads=4, MACs=223.5M, Mem (floats)=1.3M2023.12 | 12.6 | — | |
| MoA#total params=47M, Nheads=6, MACs=306.8M, Mem (floats)=1.9M2023.12 | 12.64 | — | |
| kNN-LMAdaptation=Memory, Test-time Learning=✓, + Inference Cost Compute=O(N dx), + Inference Cost Memory=307 GB, Learn Cost GPU hrs=16 hrs, Backbone=GPT2-XL (1.5B)2026.05 | 12.7 | — | |
| MoA#total params=47M, Nheads=8, MACs=390.2M, Mem (floats)=2.6M2023.12 | 12.77 | — | |
| MoA#total params=47M, Nheads=2, MACs=140.1M, Mem (floats)=0.7M2023.12 | 12.84 | — | |
| FAAST w/ readout seen WikiText103Adaptation=Fast weights, Test-time Learning=✓, + Inference Cost Compute=O(Ld2x), + Inference Cost Memory=(-99.9%), Learn Cost GPU hrs=(-93.3%), Backbone=GPT2-XL (1.5B)2026.05 | 13.23 | — | |
| GPT-NeoNumber of Parameters=1.3B2022.12 | 13.3 | — | |
| LoRa w/ same number of paramsAdaptation=Backprop, Test-time Learning=✗, Learn Cost GPU hrs=3 hrs, Backbone=GPT2-XL (1.5B)2026.05 | 13.57 | — | |
| Linear Projection (upper bound)Adaptation=Backprop, Test-time Learning=✗, + Inference Cost Compute=O(Ld2x), + Inference Cost Memory=112 MB, Learn Cost GPU hrs=3 hrs, Backbone=GPT2-XL (1.5B)2026.05 | 13.6 | — | |
| ADEPTArchitecture=GPT2 XL, Reduction=10%2026.01 | 14.7 | — | |
| FAASTAdaptation=Fast weights, Test-time Learning=✓, + Inference Cost Compute=O(Ld2x), + Inference Cost Memory=112 MB, Learn Cost GPU hrs=0.2 hrs, Backbone=GPT2-XL (1.5B)2026.05 | 15.35 | — | |
| PABEEArchitecture=GPT2 XL, Reduction=10%2026.01 | 15.4 | — | |
| Hybrid H3Number of Parameters=355M2022.12 | 16.9 | — | |
| GPT-2 XLNumber of Parameters=1.5B2022.12 | 17 | — | |
| ADEPTArchitecture=GPT2 XL, Reduction=20%2026.01 | 17 | — | |
| Pre-proj + skipModel=Pythia 410M, Params=88.1M, Evaluation protocol=Frozen probe2026.04 | 17 | — | |
| Regularized BNNorm Position=Post-Norm2022.10 | 17.1 | — | |
| Regularized BNNorm Position=Pre-Norm2022.10 | 17.1 | — | |
| BatchNormNorm Position=Post-Norm2022.10 | 17.2 | — | |
| DeeBERTArchitecture=GPT2 XL, Reduction=10%2026.01 | 17.3 | — | |
| GPT2-XL (zero-shot, lower bound)Adaptation=No adapt, Test-time Learning=✗, Backbone=GPT2-XL (1.5B)2026.05 | 17.41 | — | |
| Base ModelArchitecture=GPT2 XL, Reduction=None2026.01 | 17.5 | — | |
| DeeBERTArchitecture=GPT2 XL, Reduction=20%2026.01 | 17.5 | — | |
| KERPLE-logtrain length=2048, Extrp. length=30722022.05 | 17.56 | — | |
| PABEEArchitecture=GPT2 XL, Reduction=20%2026.01 | 17.6 | — | |
| ALiBitrain length=2048, Extrp. length=30722022.05 | 17.64 | — | |
| LoRA (r=640)Model=Pythia 410M, Params=94.4M, Evaluation protocol=Frozen probe2026.04 | 17.7 | — | |
| BatchNormNorm Position=Pre-Norm2022.10 | 17.8 | — | |
| ECPModel Size=172.12 million2024.12 | 17.82 | — | |
| KERPLE-logtrain length=2048, Extrp. length=20482022.05 | 17.84 | — | |
| ALiBitrain length=2048, Extrp. length=20482022.05 | 17.91 | — | |
| Sandwich Transformer LargeArchitecture=(s)x6 (sf)x12 (f)x6, Latency on A100 (ms)=18.92020.09 | 18.2 | — | |
| KERPLE-logtrain length=512, Extrp. length=30722022.05 | 18.24 | — | |
| KERPLE-logtrain length=512, Extrp. length=20482022.05 | 18.29 | — | |
| Full Attention2025.11 | 18.3 | — | |
| KERPLE-logtrain length=512, Extrp. length=15362022.05 | 18.37 | — | |
| Transformer-XL LargeArchitecture=(sf)x18, Latency on A100 (ms)=18.92020.09 | 18.4 | — | |
| PAR Transformer LargeArchitecture=(sfff)x7 (f)x8, Latency on A100 (ms)=13.42020.09 | 18.4 | — | |
| ALiBitrain length=512, Extrp. length=30722022.05 | 18.4 | — | |
| π-Attention2025.11 | 18.4 | — | |
| SATFORMEREvaluation Protocol=Zero-shot, Number of Parameters=1.3B, Training Tokens=30B2026.05 | 18.47 | — | |
| ALiBitrain length=512, Extrp. length=20482022.05 | 18.48 | — | |
| ALiBitrain length=512, Extrp. length=15362022.05 | 18.5 | — | |
| RESFORMEREvaluation Protocol=Zero-shot, Number of Parameters=1.3B, Training Tokens=30B2026.05 | 18.62 | — | |
| KERPLE-logtrain length=512, Extrp. length=10242022.05 | 18.76 | — | |
| ALiBitrain length=512, Extrp. length=10242022.05 | 18.81 | — | |
| TRANSFORMEREvaluation Protocol=Zero-shot, Number of Parameters=1.3B, Training Tokens=30B2026.05 | 18.83 | — | |
| Pre-proj + LoRAModel=Pythia 410M, Params=72.4M, Evaluation protocol=Frozen probe2026.04 | 19.2 | — | |
| Longformer2025.11 | 19.5 | — | |
| KERPLE-logtrain length=512, Extrp. length=5122022.05 | 19.69 | — | |
| ALiBitrain length=512, Extrp. length=5122022.05 | 19.73 | — | |
| BigBird2025.11 | 19.8 | — | |
| ECPModel Size=87 million2024.12 | 20 | — | |
| 32-bit AdamWModel=GPT2-Medium, Opt.State Mem Saved=reference, Task Setting=Fine-tuning2026.04 | 20.02 | — | |
| RingAttention2025.11 | 20.1 | — | |
| STQuantModel=GPT2-Medium, Opt.State Mem Saved=82.5%, Task Setting=Fine-tuning2026.04 | 20.1 | — | |
| GPT-2 1.5Bzero-shot=true, parameters=1.5B2023.05 | 20.13 | — | |
| Cluster-Formernumber of clusters=5122020.09 | 20.2 | — | |
| DenseModel Size=Small, Model=GPT2-S, Compression Factor (CF)=1.85x2025.12 | 20.2 | — | |
| 8-bit AdamW(bnb)Model=GPT2-Medium, Opt.State Mem Saved=75.0%, Task Setting=Fine-tuning2026.04 | 20.22 | — | |
| Cluster-Formernumber of clusters=2562020.09 | 20.3 | — | |
| Sparse Attention2020.09 | 20.5 | — | |
| Cluster-Formernumber of clusters=642020.09 | 20.5 | — |