Language Modeling on WikiText (Perplexity)
0.1PerplexitySpectrumKV
Evaluation Results
| Method | Links | |
|---|---|---|
| SpectrumKVModel=Mistral-7B, Budget (b)=0.52026.06 | 0.1 | |
| SpectrumKVModel=Gemma-2-9B, Budget (b)=0.52026.06 | 0.4 | |
| SpectrumKVModel=Qwen2.5-7B, Budget (b)=0.52026.06 | 2 | |
| Base-FP16Model=Mistral-v0.3-7B, Precision=FP162025.09 | 5.5 | |
| GPTQ-INT4Model=Mistral-v0.3-7B, Precision=4-bit2025.09 | 5.65 | |
| Base-FP16Model=Mistral-v0.3-7B-Instruct, Precision=FP162025.09 | 5.75 | |
| GPTQ-INT4Model=Mistral-v0.3-7B-Instruct, Precision=4-bit2025.09 | 5.88 | |
| FairGPTQ-INT4Model=Mistral-v0.3-7B, Precision=4-bit2025.09 | 6.28 | |
| FairGPTQ-INT4Model=Mistral-v0.3-7B-Instruct, Precision=4-bit2025.09 | 6.29 | |
| Base-FP16Model=LLaMA-3.1-8B-Instruct, Precision=FP162025.09 | 6.99 | |
| Base-FP16Model=Qwen-2.5-7B-Instruct, Precision=FP162025.09 | 7.14 | |
| GPTQ-INT4Model=LLaMA-3.1-8B-Instruct, Precision=4-bit2025.09 | 7.37 | |
| FairGPTQ-INT4Model=LLaMA-3.1-8B-Instruct, Precision=4-bit2025.09 | 7.38 | |
| GPTQ-INT4Model=Qwen-2.5-7B-Instruct, Precision=4-bit2025.09 | 7.59 | |
| FairGPTQ-INT4Model=Qwen-2.5-7B-Instruct, Precision=4-bit2025.09 | 8.26 | |
| Base-FP16Model=Qwen-3-8B, Precision=FP162025.09 | 9.52 | |
| GPTQ-INT4Model=Qwen-3-8B, Precision=4-bit2025.09 | 9.98 | |
| Base-FP16Model=OPT-6.7B, Precision=FP162025.09 | 10.24 | |
| FairGPTQ-INT4Model=Qwen-3-8B, Precision=4-bit2025.09 | 10.76 | |
| GPTQ-INT4Model=OPT-6.7B, Precision=4-bit2025.09 | 10.83 | |
| FairGPTQ-INT4Model=OPT-6.7B, Precision=4-bit2025.09 | 13.21 | |
| ROCKET-ActCostModel=Llama-3.2-1B, Compression Ratio=20%2026.06 | 14.45 | |
| ROCKET-defaultModel=Llama-3.2-1B, Compression Ratio=20%2026.06 | 14.66 | |
| CARVE + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=true, SWA Ratio=3:12026.06 | 15.41 | |
| Mamba-3 SISO + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=true, SWA Ratio=3:12026.06 | 15.54 | |
| GDN-2 + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=true, SWA Ratio=3:12026.06 | 15.62 | |
| CARVEModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=false2026.06 | 15.72 | |
| Mamba-3 MIMO + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=true, SWA Ratio=3:12026.06 | 15.81 | |
| GDN-2Model Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=false2026.06 | 15.9 | |
| Gated DeltaNet + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=true, SWA Ratio=3:12026.06 | 16 | |
| KDA + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=true, SWA Ratio=3:12026.06 | 16.01 | |
| Mamba-3 SISOModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=false2026.06 | 16.3 | |
| Gated DeltaNetModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=false2026.06 | 16.4 | |
| Mamba-3 MIMOModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=false2026.06 | 16.45 | |
| Mamba-2Model Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=false2026.06 | 16.79 | |
| KDAModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=false2026.06 | 16.81 | |
| Mamba-2 + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=true, SWA Ratio=3:12026.06 | 17.46 | |
| TransformerModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B, Sliding-Window Attention (SWA)=false, Attention Type=full attention2026.06 | 19.22 | |
| PDTrimModel=Mistral-7B, Budget (b)=0.52026.06 | 22.1 | |
| PrismScale=Large, Evaluation Protocol=Zero-shot2026.06 | 25.5 | |
| BaselineScale=Large, Evaluation Protocol=Zero-shot2026.06 | 25.7 | |
| PDTrimModel=Qwen2.5-7B, Budget (b)=0.52026.06 | 25.8 | |
| PrismScale=Medium, Evaluation Protocol=Zero-shot2026.06 | 32.9 | |
| BaselineScale=Medium, Evaluation Protocol=Zero-shot2026.06 | 33 | |
| PDTrimModel=Gemma-2-9B, Budget (b)=0.52026.06 | 35.6 | |
| PrismScale=Small, Evaluation Protocol=Zero-shot2026.06 | 51.2 | |
| BaselineScale=Small, Evaluation Protocol=Zero-shot2026.06 | 52.3 |