Language Modeling on The Pile
2.53PerplexityD3
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| D3Granularity=Ours, Architecture=LlaMa-based (1.1B), Training Tokens=100B2026.05 | 2.53 | — | |
| Dynamic LossGranularity=Sample-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B2026.05 | 2.64 | — | |
| Uniform SamplingGranularity=Sample-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B2026.05 | 2.68 | — | |
| DoReMiGranularity=Domain-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B2026.05 | 2.71 | — | |
| DoGEGranularity=Domain-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B2026.05 | 2.72 | — | |
| Data Mixing LawGranularity=Domain-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B2026.05 | 2.74 | — | |
| RegMixGranularity=Domain-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B2026.05 | 2.75 | — | |
| HGRNModel size=1b2023.11 | 4.14 | — | |
| TransformerModel size=1b2023.11 | 4.56 | — | |
| NSAModel=Qwen3 0.6B, Operation Category=Token-Level Operation, Latency@128k (Decode)=9.73, Latency@128k (Forward)=20.322026.03 | 4.57 | — | |
| DenseModel=Gemma-2-9B2026.03 | 4.63 | — | |
| Dense (full)Model=Qwen3-0.6B, Latency (128k context)=77.652026.03 | 4.66 | — | |
| Dense (full)Model=Qwen3 0.6B, Operation Category=Baseline, Latency@128k (Decode)=80.84, Latency@128k (Forward)=77.652026.03 | 4.66 | — | |
| MLAModel=Qwen3 0.6B, Operation Category=Feature-Level Operation, Latency@128k (Decode)=8.74, Latency@128k (Forward)=68.922026.03 | 4.69 | — | |
| DenseModel=Llama-3.1-8B2026.03 | 4.71 | — | |
| QuantModel=Qwen3 0.6B, Operation Category=Feature-Level Operation, Latency@128k (Decode)=72.23, Latency@128k (Forward)=59.732026.03 | 4.71 | — | |
| Wanda (25%)Model=Gemma-2-9B, Calibration Method=ZipCal (Multi-Domain)2026.03 | 4.72 | — | |
| Wanda (25%)Model=Gemma-2-9B, Calibration Method=COLA2026.03 | 4.74 | — | |
| Wanda (25%)Model=Gemma-2-9B, Calibration Method=ZipCal2026.03 | 4.76 | — | |
| AWQ (W4A16)Model=Llama-3.1-8B, Calibration Method=COLA2026.03 | 4.79 | — | |
| AWQ (W4A16)Model=Llama-3.1-8B, Calibration Method=ZipCal2026.03 | 4.81 | — | |
| GPTQ (W4A16)Model=Gemma-2-9B, Calibration Method=ZipCal2026.03 | 4.81 | — | |
| SFA (k = 16)Model=Qwen3-0.6B, Latency (128k context)=34.202026.03 | 4.81 | — | |
| SFA (k = 16)Model=Qwen3 0.6B, Operation Category=Feature-Level Operation, Latency@128k (Decode)=66.29, Latency@128k (Forward)=34.202026.03 | 4.81 | — | |
| AWQ (W4A16)Model=Llama-3.1-8B, Calibration Method=ZipCal (Multi-Domain)2026.03 | 4.82 | — | |
| GPTQ (W4A16)Model=Gemma-2-9B, Calibration Method=COLA2026.03 | 4.82 | — | |
| MLA + SFAModel=Qwen3 0.6B, Operation Category=Feature-Level Operation, Latency@128k (Decode)=6.72, Latency@128k (Forward)=65.292026.03 | 4.9 | — | |
| Wanda (25%)Model=Llama-3.1-8B, Calibration Method=ZipCal (Multi-Domain)2026.03 | 4.94 | — | |
| +SFA (k = 16)Model=Qwen3 0.6B, Operation Category=Token-Level Operation, Latency@128k (Decode)=8.85, Latency@128k (Forward)=17.172026.03 | 4.95 | — | |
| Wanda (25%)Model=Llama-3.1-8B, Calibration Method=COLA2026.03 | 4.98 | — | |
| Wanda (25%)Model=Llama-3.1-8B, Calibration Method=ZipCal2026.03 | 4.98 | — | |
| GPTQ (W4A16)Model=Gemma-2-9B, Calibration Method=ZipCal (Multi-Domain)2026.03 | 4.99 | — | |
| LRUModel size=1b2023.11 | 5.07 | — | |
| GPTQ (W4A16)Model=Llama-3.1-8B, Calibration Method=COLA2026.03 | 5.13 | — | |
| GPTQ (W4A16)Model=Llama-3.1-8B, Calibration Method=ZipCal2026.03 | 5.15 | — | |
| SFA (quant)Model=Qwen3 0.6B, Operation Category=Feature-Level Operation, Latency@128k (Decode)=57.47, Latency@128k (Forward)=30.742026.03 | 5.16 | — | |
| GPTQ (W4A16)Model=Llama-3.1-8B, Calibration Method=ZipCal (Multi-Domain)2026.03 | 5.34 | — | |
| Hybrid H3Number of Parameters=2.7B2022.12 | 5.4 | — | |
| Low-RankModel=Qwen3 0.6B, Operation Category=Feature-Level Operation, Latency@128k (Decode)=40.58, Latency@128k (Forward)=32.462026.03 | 5.5 | — | |
| GPT-NeoNumber of Parameters=2.7B2022.12 | 5.7 | — | |
| Hybrid H3Number of Parameters=1.3B2022.12 | 6 | — | |
| Dense (d = 64)Model=Qwen3-0.6B, Latency (128k context)=30.842026.03 | 6.03 | — | |
| Short (d = 64)Model=Qwen3 0.6B, Operation Category=Feature-Level Operation, Latency@128k (Decode)=38.68, Latency@128k (Forward)=30.842026.03 | 6.03 | — | |
| MeSHModel Backbone=Pythia-6.9B, Scheme=Recursive (-31.25%), Layers=6+10R2+6, Variant=+mesh2025.10 | 6.09 | — | |
| GPT-NeoNumber of Parameters=1.3B2022.12 | 6.2 | — | |
| Recursive baselineModel Backbone=Pythia-6.9B, Scheme=Recursive (-31.25%), Layers=6+10R2+6, Variant=base2025.10 | 6.29 | — | |
| Baseline (no intervention)Model=Qwen2.5-32B-Instruct (64 layers)2026.03 | 6.41 | — | |
| RFAModel=Qwen2.5-32B-Instruct (64 layers)2026.03 | 6.44 | — | |
| 2SSP (25%)Model=Gemma-2-9B, Calibration Method=COLA2026.03 | 6.47 | — | |
| 2SSP (25%)Model=Gemma-2-9B, Calibration Method=ZipCal2026.03 | 6.47 | — | |
| AcTModel=Qwen2.5-32B-Instruct (64 layers)2026.03 | 6.49 | — | |
| PCA-OT1Model=Qwen2.5-32B-Instruct (64 layers)2026.03 | 6.55 | — | |
| Baseline (no intervention)Model=Qwen2.5-14B-Instruct (48 layers)2026.03 | 6.63 | — | |
| MeSHModel Backbone=Pythia-2.8B, Scheme=Recursive (-31.25%), Layers=6+10R2+6, Variant=+mesh2025.10 | 6.7 | — | |
| RFAModel=Qwen2.5-14B-Instruct (48 layers)2026.03 | 6.72 | — | |
| 2SSP (25%)Model=Gemma-2-9B, Calibration Method=ZipCal (Multi-Domain)2026.03 | 6.86 | — | |
| Recursive baselineModel Backbone=Pythia-2.8B, Scheme=Recursive (-31.25%), Layers=6+10R2+6, Variant=base2025.10 | 6.9 | — | |
| PCA-OT2Model=Qwen2.5-32B-Instruct (64 layers)2026.03 | 6.99 | — | |
| AcTModel=Qwen2.5-14B-Instruct (48 layers)2026.03 | 7.02 | — | |
| Hybrid H3Number of Parameters=355M2022.12 | 7.1 | — | |
| MeSHModel Size=Pythia-1.4B, Scheme=Recursive, Layers=4+8R2+4, Variant=+mesh, Parameter reduction percentage=-33.3%2025.10 | 7.39 | — | |
| VanillaModel Size=Pythia-1.4B, Scheme=Vanilla, Layers=242025.10 | 7.44 | — | |
| RecursiveModel Size=Pythia-1.4B, Scheme=Recursive, Layers=4+8R2+4, Variant=+anchor*, Parameter reduction percentage=-33.3%2025.10 | 7.51 | — | |
| RecursiveModel Size=Pythia-1.4B, Scheme=Recursive, Layers=4+8R2+4, Variant=+anchor, Parameter reduction percentage=-33.3%2025.10 | 7.51 | — | |
| PCA-OT2Model=Qwen2.5-14B-Instruct (48 layers)2026.03 | 7.54 | — | |
| Baseline (no intervention)Model=Qwen2.5-7B-Instruct (28 layers)2026.03 | 7.55 | — | |
| RecursiveModel Size=Pythia-1.4B, Scheme=Recursive, Layers=4+8R2+4, Variant=+residual, Parameter reduction percentage=-33.3%2025.10 | 7.58 | — | |
| RecursiveModel Size=Pythia-1.4B, Scheme=Recursive, Layers=4+8R2+4, Variant=base, Parameter reduction percentage=-33.3%2025.10 | 7.63 | — | |
| AcTModel=Qwen2.5-7B-Instruct (28 layers)2026.03 | 7.85 | — | |
| MeSHModel Size=Pythia-1B, Scheme=Recursive, Layers=3+5R2+3, Variant=+mesh, Parameter reduction percentage=-31.3%2025.10 | 7.9 | — | |
| RFAModel=Qwen2.5-7B-Instruct (28 layers)2026.03 | 7.94 | — | |
| VanillaModel Size=Pythia-1B, Scheme=Vanilla, Layers=162025.10 | 7.96 | — | |
| Baseline (no intervention)Model=Llama-2-13b-chat-hf (40 layers)2026.03 | 8.01 | — | |
| RFAModel=Llama-2-13b-chat-hf (40 layers)2026.03 | 8.04 | — | |
| RecursiveModel Size=Pythia-1B, Scheme=Recursive, Layers=3+5R2+3, Variant=+anchor*, Parameter reduction percentage=-31.3%2025.10 | 8.07 | — | |
| RecursiveModel Size=Pythia-1B, Scheme=Recursive, Layers=3+5R2+3, Variant=+anchor, Parameter reduction percentage=-31.3%2025.10 | 8.1 | — | |
| PCA-OT1Model=Qwen2.5-14B-Instruct (48 layers)2026.03 | 8.16 | — | |
| RecursiveModel Size=Pythia-1B, Scheme=Recursive, Layers=3+5R2+3, Variant=+residual, Parameter reduction percentage=-31.3%2025.10 | 8.19 | — | |
| RecursiveModel Size=Pythia-1B, Scheme=Recursive, Layers=3+5R2+3, Variant=base, Parameter reduction percentage=-31.3%2025.10 | 8.2 | — | |
| PCA-OT1Model=Llama-2-13b-chat-hf (40 layers)2026.03 | 8.41 | — | |
| ContextLMModel Size=410M2026.02 | 8.67 | 15.82 | |
| Baseline (no intervention)Model=Llama-3.1-8B-Instruct (32 layers)2026.03 | 8.68 | — | |
| ConceptLMModel Size=410M2026.02 | 8.7 | 15.31 | |
| RFAModel=Llama-3.1-8B-Instruct (32 layers)2026.03 | 8.75 | — | |
| Hybrid H3Number of Parameters=125M2022.12 | 8.8 | — | |
| Pythia-410MModel Size=410M2026.02 | 8.88 | 17.84 | |
| 2SSP (25%)Model=Llama-3.1-8B, Calibration Method=ZipCal2026.03 | 9 | — | |
| VanillaModel Size=Pythia-410M, Scheme=Vanilla, Layers=242025.10 | 9.07 | — | |
| MeSHModel Size=Pythia-410M, Scheme=Recursive, Layers=4+8R2+4, Variant=+mesh, Parameter reduction percentage=-33.3%2025.10 | 9.09 | — | |
| RFAModel=Llama-2-7b-chat-hf (32 layers)2026.03 | 9.16 | — | |
| PCA-OT1Model=Llama-3.1-8B-Instruct (32 layers)2026.03 | 9.18 | — | |
| RecursiveModel Size=Pythia-410M, Scheme=Recursive, Layers=4+8R2+4, Variant=+anchor, Parameter reduction percentage=-33.3%2025.10 | 9.19 | — | |
| Baseline (no intervention)Model=Llama-2-7b-chat-hf (32 layers)2026.03 | 9.21 | — | |
| RecursiveModel Size=Pythia-410M, Scheme=Recursive, Layers=4+8R2+4, Variant=base, Parameter reduction percentage=-33.3%2025.10 | 9.31 | — | |
| MeSHModel Size=Pythia-410M, Scheme=Recursive, Layers=3+6R3+3, Variant=+mesh, Parameter reduction percentage=-50.0%2025.10 | 9.35 | — | |
| PCA-OT1Model=Llama-2-7b-chat-hf (32 layers)2026.03 | 9.36 | — | |
| 2SSP (25%)Model=Llama-3.1-8B, Calibration Method=COLA2026.03 | 9.36 | — | |
| GPT-NeoNumber of Parameters=125M2022.12 | 9.4 | — | |
| PCA-OT2Model=Llama-2-13b-chat-hf (40 layers)2026.03 | 9.4 | — | |
| PCA-OT1Model=Qwen2.5-7B-Instruct (28 layers)2026.03 | 9.4 | — |