Language Modeling on C4 (Perplexity, Memory)
9.44PerplexityMeta-Llama-3-8B Dense
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Meta-Llama-3-8B Dense2026.03 | 9.44 | — | |
| 3BASiL-TMConfig=4:8+64LR, One-shot=true2026.03 | 13.02 | — | |
| 3BASiLConfig=4:8+64LR, One-shot=true2026.03 | 13.74 | — | |
| Hassle-free-ALPSConfig=4:8+64LR, One-shot=true2026.03 | 14.04 | — | |
| 3BASiL-TMConfig=2:4+64LR, One-shot=true2026.03 | 14.34 | — | |
| Hassle-free-SparseGPTConfig=4:8+64LR, One-shot=true2026.03 | 14.65 | — | |
| PAMMModel size=1B, Rank (r)=1/128, Training tokens=13.1B2025.06 | 15.01 | 78 | |
| PAMMModel size=1B, Rank (r)=1/256, Training tokens=13.1B2025.06 | 15.06 | 42 | |
| PAMMModel size=1B, Rank (r)=1/512, Training tokens=13.1B2025.06 | 15.36 | 24 | |
| Full RankModel size=1B, Training tokens=13.1B2025.06 | 15.56 | 3 | |
| 3BASiLConfig=2:4+64LR, One-shot=true2026.03 | 15.76 | — | |
| Hassle-free-ALPSConfig=2:4+64LR, One-shot=true2026.03 | 16.15 | — | |
| OATSConfig=4:8+64LR, One-shot=true2026.03 | 16.38 | — | |
| Hassle-free-SparseGPTConfig=2:4+64LR, One-shot=true2026.03 | 17.77 | — | |
| 3BASiL-TMConfig=3:8+64LR, One-shot=true2026.03 | 18.11 | — | |
| PAMMModel size=350M, Rank (r)=1/128, Training tokens=6.4B2025.06 | 18.4 | 42 | |
| PAMMModel size=350M, Rank (r)=1/256, Training tokens=6.4B2025.06 | 18.48 | 24 | |
| PAMMModel size=350M, Rank (r)=1/512, Training tokens=6.4B2025.06 | 18.49 | 15 | |
| Full RankModel size=350M, Training tokens=6.4B2025.06 | 18.8 | 1.5 | |
| OATSConfig=2:4+64LR, One-shot=true2026.03 | 21.59 | — | |
| 3BASiLConfig=3:8+64LR, One-shot=true2026.03 | 23.07 | — | |
| Hassle-free-ALPSConfig=3:8+64LR, One-shot=true2026.03 | 23.93 | — | |
| Hassle-free-SparseGPTConfig=3:8+64LR, One-shot=true2026.03 | 29.32 | — | |
| Full RankModel size=60M, Training tokens=1.1B2025.06 | 30.97 | 256 | |
| PAMMModel size=60M, Rank (r)=1/128, Training tokens=1.1B2025.06 | 31.94 | 8 | |
| PAMMModel size=60M, Rank (r)=1/256, Training tokens=1.1B2025.06 | 32.18 | 5 | |
| PAMMModel size=60M, Rank (r)=1/512, Training tokens=1.1B2025.06 | 32.53 | 3.5 | |
| OATSConfig=3:8+64LR, One-shot=true2026.03 | 58.88 | — |