Language Modeling on FineWeb-Edu
8.32PPLBase Model
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Base ModelModel=Qwen3-32B, Sparsity=0.02026.05 | 8.32 | — | |
| Base ModelModel=Llama-3.1-8B, Sparsity=0.02026.05 | 8.375 | — | |
| Looped Hybrid (Full+GDN)Parameters=1.3B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=true, Architecture Type=Hybrid LT22026.05 | 9.12 | — | |
| Base ModelModel=Qwen3-14B, Sparsity=0.02026.05 | 9.43 | — | |
| Looped Hybrid (GDN+DSA)Parameters=1.3B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=true, Architecture Type=Hybrid LT22026.05 | 9.5 | — | |
| Looped KDAParameters=1.3B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=true, Architecture Type=LT2-linear attention2026.05 | 9.68 | — | |
| Looped GDNParameters=1.3B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=true, Architecture Type=LT2-linear attention2026.05 | 9.75 | — | |
| Looped Hybrid (Full+DSA)Parameters=1.3B, Token Budget=100B, Loops (K)=4, Architecture Type=Hybrid LT22026.05 | 9.8 | — | |
| Looped Hybrid (Full+Window)Parameters=1.3B, Token Budget=100B, Loops (K)=4, Architecture Type=Hybrid LT22026.05 | 9.84 | — | |
| Looped Transformer (ref)Parameters=1.3B, Token Budget=100B, Loops (K)=4, Architecture Type=Baseline2026.05 | 9.87 | — | |
| Looped DSAParameters=1.3B, Token Budget=100B, Loops (K)=4, Architecture Type=LT2-sparse attention2026.05 | 9.97 | — | |
| Looped NSAParameters=1.3B, Token Budget=100B, Loops (K)=4, Architecture Type=LT2-sparse attention2026.05 | 10.17 | — | |
| Mamba-3-MIMO-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B, MIMO Rank=42026.03 | 10.24 | — | |
| Looped Mamba2Parameters=1.3B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=false, Architecture Type=LT2-linear attention2026.05 | 10.3 | — | |
| Mamba-3-SISO-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 10.35 | — | |
| Looped WindowParameters=1.3B, Token Budget=100B, Loops (K)=4, Architecture Type=LT2-sparse attention2026.05 | 10.42 | — | |
| GDN-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 10.45 | — | |
| Mamba-2-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 10.47 | — | |
| Transformer-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 10.51 | — | |
| GPTParameters=1.3B, Training tokens=100B2025.11 | 10.52 | — | |
| TransformerParameters=1.3B, Token Budget=100B, Loops (K)=4, Architecture Type=Baseline2026.05 | 10.65 | — | |
| NextLat (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 10.83 | — | |
| NextLat (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 10.88 | — | |
| MTP (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 10.9 | — | |
| FP16 (Baseline)Bitwidth=16 bits2026.05 | 10.96 | — | |
| MTP (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 11 | — | |
| JTP (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 11.08 | — | |
| Mamba-3-MIMO-880MTraining Tokens=100B, Context Length=2K, Scale=880M, MIMO Rank=42026.03 | 11.11 | — | |
| RCOg=32, s=50, Bitwidth=4.0 bits, Wall-clock time (Wall)=117 m2026.05 | 11.14 | — | |
| RCOg=16, s=200, Bitwidth=4.0 bits, Wall-clock time (Wall)=215 m2026.05 | 11.16 | — | |
| EvoPressgenerations=100, Bitwidth=4.0 bits, Wall-clock time (Wall)=11–14 h2026.05 | 11.16 | — | |
| RCOg=4, s=200, Bitwidth=4.0 bits, Wall-clock time (Wall)=62 m2026.05 | 11.17 | — | |
| JTP (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 11.18 | — | |
| Mamba-3-SISO-880MTraining Tokens=100B, Context Length=2K, Scale=880M2026.03 | 11.23 | — | |
| HIGGSQuantization Method/Protocol=Linear surrogate, DP, Bitwidth=4.0 bits, Wall-clock time (Wall)=∼2 h2026.05 | 11.28 | — | |
| Mamba-2-880MTraining Tokens=100B, Context Length=2K, Scale=880M2026.03 | 11.35 | — | |
| GDN-880MTraining Tokens=100B, Context Length=2K, Scale=880M2026.03 | 11.37 | — | |
| RCOg=4, s=200, Bitwidth=3.5 bits, Wall-clock time (Wall)=62 m2026.05 | 11.41 | — | |
| Transformer-880MTraining Tokens=100B, Context Length=2K, Scale=880M2026.03 | 11.42 | — | |
| IMPQQuantization Method/Protocol=Binary bitwidth; Shapley surrogate, MILP, Bitwidth=4.0 bits, Wall-clock time (Wall)=∼2 h2026.05 | 11.42 | — | |
| Looped Hybrid (Full+GDN)Parameters=0.6B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=true, Architecture Type=Hybrid LT22026.05 | 11.43 | — | |
| RCOg=16, s=200, Bitwidth=3.5 bits, Wall-clock time (Wall)=215 m2026.05 | 11.45 | — | |
| RCOg=32, s=50, Bitwidth=3.5 bits, Wall-clock time (Wall)=117 m2026.05 | 11.46 | — | |
| EvoPressgenerations=100, Bitwidth=3.5 bits, Wall-clock time (Wall)=11–14 h2026.05 | 11.5 | — | |
| HIGGSQuantization Method/Protocol=Linear surrogate, DP, Bitwidth=3.5 bits, Wall-clock time (Wall)=∼2 h2026.05 | 11.65 | — | |
| Looped Hybrid (GDN+DSA)Parameters=0.6B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=true, Architecture Type=Hybrid LT22026.05 | 11.85 | — | |
| Looped Transformer (ref)Parameters=0.6B, Token Budget=100B, Loops (K)=4, Architecture Type=Baseline2026.05 | 11.92 | — | |
| Looped GDNParameters=0.6B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=true, Architecture Type=LT2-linear attention2026.05 | 12.06 | — | |
| Looped DSAParameters=0.6B, Token Budget=100B, Loops (K)=4, Architecture Type=LT2-sparse attention2026.05 | 12.08 | — | |
| Looped KDAParameters=0.6B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=true, Architecture Type=LT2-linear attention2026.05 | 12.13 | — | |
| Looped Hybrid (Full+DSA)Parameters=0.6B, Token Budget=100B, Loops (K)=4, Architecture Type=Hybrid LT22026.05 | 12.2 | — | |
| Looped Hybrid (Full+Window)Parameters=0.6B, Token Budget=100B, Loops (K)=4, Architecture Type=Hybrid LT22026.05 | 12.24 | — | |
| Looped NSAParameters=0.6B, Token Budget=100B, Loops (K)=4, Architecture Type=LT2-sparse attention2026.05 | 12.3 | — | |
| IMPQQuantization Method/Protocol=Binary bitwidth; Shapley surrogate, MILP, Bitwidth=3.5 bits, Wall-clock time (Wall)=∼2 h2026.05 | 12.31 | — | |
| Elastic Memory_expMem. Size=0.8M, Add. Params=02026.02 | 12.318 | 11.932 | |
| Elastic Memory_uniMem. Size=0.8M, Add. Params=02026.02 | 12.419 | 10.915 | |
| MelodiMem. Size=0.8M, Add. Params=1.5M2026.02 | 12.509 | 23.532 | |
| Mamba-3-MIMO-440MTraining Tokens=100B, Context Length=2K, Scale=440M, MIMO Rank=42026.03 | 12.72 | — | |
| Infini-TransformerMem. Size=0.8M, Add. Params=962026.02 | 12.736 | 24.417 | |
| Looped Mamba2Parameters=0.6B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=false, Architecture Type=LT2-linear attention2026.05 | 12.78 | — | |
| Mamba-3-SISO-440MTraining Tokens=100B, Context Length=2K, Scale=440M2026.03 | 12.87 | — | |
| Looped WindowParameters=0.6B, Token Budget=100B, Loops (K)=4, Architecture Type=LT2-sparse attention2026.05 | 12.87 | — | |
| Transformer++Mem. Size=0, Add. Params=02026.02 | 12.897 | 29.882 | |
| Mem. TransformerMem. Size=0.8M, Add. Params=962026.02 | 12.924 | 29.963 | |
| Mamba-2-440MTraining Tokens=100B, Context Length=2K, Scale=440M2026.03 | 13 | — | |
| GDN-440MTraining Tokens=100B, Context Length=2K, Scale=440M2026.03 | 13.01 | — | |
| Transformer-440MTraining Tokens=100B, Context Length=2K, Scale=440M2026.03 | 13.03 | — | |
| TransformerParameters=0.6B, Token Budget=100B, Loops (K)=4, Architecture Type=Baseline2026.05 | 13.14 | — | |
| Looped DeltaNetParameters=0.6B, Token Budget=100B, Loops (K)=4, D-Gate=false, Δ=true, Architecture Type=LT2-linear attention2026.05 | 14.16 | — | |
| Looped HGRN2Parameters=0.6B, Token Budget=100B, Loops (K)=4, D-Gate=true, Δ=false, Architecture Type=LT2-linear attention2026.05 | 14.59 | — | |
| EvoPressgenerations=100, Bitwidth=2.5 bits, Wall-clock time (Wall)=11–14 h2026.05 | 15.33 | — | |
| RCOg=32, s=50, Bitwidth=2.5 bits, Wall-clock time (Wall)=117 m2026.05 | 15.43 | — | |
| RCOg=16, s=200, Bitwidth=2.5 bits, Wall-clock time (Wall)=215 m2026.05 | 15.45 | — | |
| RCOg=4, s=200, Bitwidth=2.5 bits, Wall-clock time (Wall)=62 m2026.05 | 15.78 | — | |
| Mamba-3-MIMO-180MTraining Tokens=100B, Context Length=2K, Scale=180M, MIMO Rank=42026.03 | 16.46 | — | |
| GDN-180MTraining Tokens=100B, Context Length=2K, Scale=180M2026.03 | 16.52 | — | |
| Mamba-3-SISO-180MTraining Tokens=100B, Context Length=2K, Scale=180M2026.03 | 16.59 | — | |
| Mamba-2-180MTraining Tokens=100B, Context Length=2K, Scale=180M2026.03 | 16.76 | — | |
| Transformer-180MTraining Tokens=100B, Context Length=2K, Scale=180M2026.03 | 16.89 | — | |
| PutriModel=Qwen3-32B, Sparsity=0.52026.05 | 16.95 | — | |
| IMPQQuantization Method/Protocol=Binary bitwidth; Shapley surrogate, MILP, Bitwidth=2.5 bits, Wall-clock time (Wall)=∼2 h2026.05 | 18.79 | — | |
| PutriModel=Qwen3-14B, Sparsity=0.52026.05 | 19.97 | — | |
| HIGGSQuantization Method/Protocol=Linear surrogate, DP, Bitwidth=2.5 bits, Wall-clock time (Wall)=∼2 h2026.05 | 20.08 | — | |
| RCOg=16, s=200, Bitwidth=2.25 bits, Wall-clock time (Wall)=215 m2026.05 | 20.47 | — | |
| RCOg=4, s=200, Bitwidth=2.25 bits, Wall-clock time (Wall)=62 m2026.05 | 20.6 | — | |
| RCOg=32, s=50, Bitwidth=2.25 bits, Wall-clock time (Wall)=117 m2026.05 | 20.66 | — | |
| EvoPressgenerations=100, Bitwidth=2.25 bits, Wall-clock time (Wall)=11–14 h2026.05 | 21.4 | — | |
| 2SSPModel=Qwen3-32B, Sparsity=0.52026.05 | 21.67 | — | |
| IMPQQuantization Method/Protocol=Binary bitwidth; Shapley surrogate, MILP, Bitwidth=2.25 bits, Wall-clock time (Wall)=∼2 h2026.05 | 24.18 | — | |
| 2SSPModel=Qwen3-14B, Sparsity=0.52026.05 | 25 | — | |
| HIGGSQuantization Method/Protocol=Linear surrogate, DP, Bitwidth=2.25 bits, Wall-clock time (Wall)=∼2 h2026.05 | 32.07 | — | |
| BlockPrunerModel=Qwen3-32B, Sparsity=0.52026.05 | 35.03 | — | |
| 2SSPModel=Llama-3.1-8B, Sparsity=0.52026.05 | 41.25 | — | |
| PutriModel=Llama-3.1-8B, Sparsity=0.52026.05 | 41.75 | — | |
| PutriModel=Qwen3-32B, Sparsity=0.752026.05 | 41.78 | — | |
| PutriModel=Qwen3-14B, Sparsity=0.752026.05 | 68.75 | — | |
| EvoPressModel=Qwen3-14B, Sparsity=0.52026.05 | 73.75 | — | |
| EvoPressModel=Llama-3.1-8B, Sparsity=0.52026.05 | 96 | — | |
| 2SSPModel=Qwen3-32B, Sparsity=0.752026.05 | 99.62 | — | |
| BlockPrunerModel=Qwen3-14B, Sparsity=0.52026.05 | 131.5 | — |