Language Modeling on Lambada (OpenAI Split)
3.11PPLLlama 3 8B Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Llama 3 8B InstructBackbone=Llama 3 8B Instruct, Pruning ratio=0%, Linear transform=None2025.05 | 3.11 | — | |
| MeSHModel Size=Pythia-1.4B, Scheme=Recursive, Layers=4+8R2+4, Variant=+mesh, Parameter reduction percentage=-33.3%2025.10 | 9.72 | — | |
| SPIRALFORMER-LModel Scale=1.4B, Topology=MeSH2026.02 | 9.73 | — | |
| VanillaModel Size=Pythia-1.4B, Scheme=Vanilla, Layers=242025.10 | 10.51 | — | |
| MeSHModel Size=Pythia-1B, Scheme=Recursive, Layers=3+5R2+3, Variant=+mesh, Parameter reduction percentage=-31.3%2025.10 | 12.19 | — | |
| LLMPrunerBackbone=Llama 3 8B Instruct, Pruning ratio=25%, Linear transform=None2025.05 | 12.31 | — | |
| VanillaModel Size=Pythia-1B, Scheme=Vanilla, Layers=162025.10 | 13.53 | — | |
| ReplaceMeBackbone=Llama 3 8B Instruct, Pruning ratio=25%, Linear transform=Multi_LT_NC (Cosine)2025.05 | 13.95 | — | |
| ReplaceMeBackbone=Llama 3 8B Instruct, Pruning ratio=25%, Linear transform=Linear (Cosine)2025.05 | 15.88 | — | |
| VanillaModel Size=Pythia-410M, Scheme=Vanilla, Layers=242025.10 | 19.48 | — | |
| MeSHModel Size=Pythia-410M, Scheme=Recursive, Layers=4+8R2+4, Variant=+mesh, Parameter reduction percentage=-33.3%2025.10 | 19.63 | — | |
| ConceptLMModel Size=160M2026.02 | 20.01 | 34.23 | |
| ReplaceMeBackbone=Llama 3 8B Instruct, Pruning ratio=25%, Linear transform=Linear (LS)2025.05 | 20.23 | — | |
| MeSHModel Size=Pythia-410M, Scheme=Recursive, Layers=3+6R3+3, Variant=+mesh, Parameter reduction percentage=-50.0%2025.10 | 20.72 | — | |
| ContextLMModel Size=160M2026.02 | 25.97 | 43.09 | |
| SVD-LLMBackbone=Llama 3 8B Instruct, Pruning ratio=25%, Linear transform=None2025.05 | 29.9 | — | |
| Pythia-160MModel Size=160M2026.02 | 38.2 | 67.87 | |
| VanillaModel Size=160M, Scheme=Vanilla, Layers=122025.10 | 42.86 | — | |
| MeSHModel Size=160M, Scheme=Recursive, Layers=2+4R2+2, Variant=+mesh, Parameter reduction percentage=-33.3%2025.10 | 46.6 | — | |
| ConceptLMModel Size=124M2026.02 | 51.83 | 84.75 | |
| GPT2-124MModel Size=124M2026.02 | 69.29 | 101.94 | |
| UIDLBackbone=Llama 3 8B Instruct, Pruning ratio=25%, Linear transform=Identity2025.05 | 2,216.96 | — |