Language Modeling on OpenWebText (val)
2.6091Validation LossLeapfrog
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LeapfrogModel Size=GPT Large, Number of Parameters=772M2025.11 | 2.6091 | — | — | — | — | — | — | — | — | — | — | |
| MidpointModel Size=GPT Large, Number of Parameters=772M2025.11 | 2.6183 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 XLargeLayers=24, Initialization=rand init, Steps=200K2024.04 | 2.62 | — | — | — | — | — | — | — | — | — | — | |
| Inheritune (GPT-2 XLarge)Layers=24, Initialization=Inheritune, Steps=100K2024.04 | 2.64 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 XLargeLayers=48, Initialization=rand init, Steps=100K2024.04 | 2.65 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 XLargeLayers=24, Initialization=rand init, Steps=100K2024.04 | 2.69 | — | — | — | — | — | — | — | — | — | — | |
| BaselineModel Size=GPT Large, Number of Parameters=772M2025.11 | 2.6988 | — | — | — | — | — | — | — | — | — | — | |
| ADAMSsynchronization=DENSE all-reduce2026.07 | 2.73 | — | — | — | — | — | — | — | — | — | — | |
| SCAPEd=0.12026.07 | 2.73 | — | — | — | — | — | — | — | — | — | — | |
| ADAMWsynchronization=DENSE all-reduce2026.07 | 2.76 | — | — | — | — | — | — | — | — | — | — | |
| SCAPEd=0.012026.07 | 2.76 | — | — | — | — | — | — | — | — | — | — | |
| Inheritune (GPT-2 Large)Layers=18, Initialization=Inheritune, Steps=100K2024.04 | 2.8 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 MediumLayers=24, Initialization=rand init, Steps=100K2024.04 | 2.81 | — | — | — | — | — | — | — | — | — | — | |
| Inheritune (GPT-2 Medium)Layers=16, Initialization=Inheritune, Steps=100K, Note=Final Model2024.04 | 2.81 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 MediumLayers=16, Initialization=rand init, Steps=200K2024.04 | 2.83 | — | — | — | — | — | — | — | — | — | — | |
| Inheritune (GPT-2 Medium)Layers=14, Initialization=Inheritune, Steps=100K2024.04 | 2.84 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 LargeLayers=18, Initialization=rand init, Steps=200K2024.04 | 2.84 | — | — | — | — | — | — | — | — | — | — | |
| GatedFWA + NSANumber of Parameters=362.7M, Context Length=40962025.12 | 2.842 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 LargeLayers=36, Initialization=rand init, Steps=100K2024.04 | 2.85 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2-baseModel size=124M2024.05 | 2.85 | — | — | — | — | — | — | — | — | — | — | |
| GatedFWA + NSANumber of Parameters=362.7M, Context Length=10242025.12 | 2.859 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 MediumLayers=16, Initialization=rand init, Steps=100K2024.04 | 2.86 | — | — | — | — | — | — | — | — | — | — | |
| MambaNumber of Parameters=371.5M, Context Length=40962025.12 | 2.868 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA) + SWA + NSANumber of Parameters=361.8M, Context Length=40962025.12 | 2.868 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA) + SWA + NSANumber of Parameters=361.8M, Context Length=10242025.12 | 2.87 | — | — | — | — | — | — | — | — | — | — | |
| Inheritune (GPT-2 Medium)Layers=12, Initialization=Inheritune, Steps=100K2024.04 | 2.87 | — | — | — | — | — | — | — | — | — | — | |
| GatedFWANumber of Parameters=360.7M, Context Length=40962025.12 | 2.871 | — | — | — | — | — | — | — | — | — | — | |
| GatedFWANumber of Parameters=360.7M, Context Length=10242025.12 | 2.874 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA)Number of Parameters=357.7M, Context Length=40962025.12 | 2.883 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA) + SWANumber of Parameters=357.7M, Context Length=40962025.12 | 2.887 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA)Number of Parameters=357.7M, Context Length=10242025.12 | 2.891 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA) + SWANumber of Parameters=357.7M, Context Length=10242025.12 | 2.892 | — | — | — | — | — | — | — | — | — | — | |
| MambaNumber of Parameters=371.5M, Context Length=10242025.12 | 2.902 | — | — | — | — | — | — | — | — | — | — | |
| BaselineModel Size=GPT Small, Number of Parameters=124M2025.11 | 2.9022 | — | — | — | — | — | — | — | — | — | — | |
| MidpointModel Size=GPT Small, Number of Parameters=124M2025.11 | 2.9148 | — | — | — | — | — | — | — | — | — | — | |
| RWKVNumber of Parameters=354.8M, Context Length=40962025.12 | 2.931 | — | — | — | — | — | — | — | — | — | — | |
| TMMFormer2026.05 | 2.9342 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2normalization=LayerNorm (LN)2025.12 | 2.94 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2normalization=Derf2025.12 | 2.94 | — | — | — | — | — | — | — | — | — | — | |
| LNArchitecture=GPT-2 (124M)2025.12 | 2.94 | — | — | — | — | — | — | — | — | — | — | |
| DerfArchitecture=GPT-2 (124M)2025.12 | 2.94 | — | — | — | — | — | 0 | 0.03 | — | — | — | |
| LeapfrogModel Size=GPT Small, Number of Parameters=124M2025.11 | 2.9432 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2normalization=RMSNorm2025.12 | 2.95 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2normalization=DyT2025.12 | 2.97 | — | — | — | — | — | — | — | — | — | — | |
| DyTArchitecture=GPT-2 (124M)2025.12 | 2.97 | — | — | — | — | — | — | — | — | — | — | |
| GPT-2 LargeLayers=18, Initialization=rand init, Steps=100K2024.04 | 2.97 | — | — | — | — | — | — | — | — | — | — | |
| RWKVNumber of Parameters=354.8M, Context Length=10242025.12 | 2.983 | — | — | — | — | — | — | — | — | — | — | |
| AdamWFormer2026.05 | 2.9883 | — | — | — | — | — | — | — | — | — | — | |
| AdamFormer2026.05 | 2.9911 | — | — | — | — | — | — | — | — | — | — | |
| GLANumber of Parameters=361.1M, Context Length=40962025.12 | 3.001 | — | — | — | — | — | — | — | — | — | — | |
| VanillaTransformer2026.05 | 3.0078 | — | — | — | — | — | — | — | — | — | — | |
| MuonFormer2026.05 | 3.0096 | — | — | — | — | — | — | — | — | — | — | |
| sHCModel Scale=L2026.03 | 3.012 | — | — | — | — | — | — | — | — | — | — | |
| GLANumber of Parameters=361.1M, Context Length=10242025.12 | 3.018 | — | — | — | — | — | — | — | — | — | — | |
| mHCModel Scale=L2026.03 | 3.023 | — | — | — | — | — | — | — | — | — | — | |
| mHC-liteModel Scale=L2026.03 | 3.023 | — | — | — | — | — | — | — | — | — | — | |
| RCModel Scale=L2026.03 | 3.066 | — | — | — | — | — | — | — | — | — | — | |
| GPT2-81M-LOOPParameters=81M, Layers=6, Loops=62024.09 | 3.11 | — | — | — | — | — | — | — | — | — | — | |
| GPT2-124MParameters=124M, Layers=12, Loops=12024.09 | 3.12 | — | — | — | — | — | — | — | — | — | — | |
| HCModel Scale=L2026.03 | 3.132 | — | — | — | — | — | — | — | — | — | — | |
| CRATE-α-baseModel size=120M2024.05 | 3.14 | — | — | — | — | — | — | — | — | — | — | |
| GPT2-67M-LOOPParameters=67M, Layers=4, Loops=122024.09 | 3.15 | — | — | — | — | — | — | — | — | — | — | |
| GatedFWA + NSANumber of Parameters=126.1M, Context Length=10242025.12 | 3.215 | — | — | — | — | — | — | — | — | — | — | |
| RetNetNumber of Parameters=373.2M, Context Length=40962025.12 | 3.227 | — | — | — | — | — | — | — | — | — | — | |
| GatedFWA + NSANumber of Parameters=126.1M, Context Length=40962025.12 | 3.23 | — | — | — | — | — | — | — | — | — | — | |
| MambaNumber of Parameters=129.2M, Context Length=40962025.12 | 3.231 | — | — | — | — | — | — | — | — | — | — | |
| GatedFWANumber of Parameters=125.1M, Context Length=10242025.12 | 3.237 | — | — | — | — | — | — | — | — | — | — | |
| MambaNumber of Parameters=129.2M, Context Length=10242025.12 | 3.238 | — | — | — | — | — | — | — | — | — | — | |
| sHCModel Scale=M2026.03 | 3.239 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA) + SWA + NSANumber of Parameters=125.4M, Context Length=10242025.12 | 3.24 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA)Number of Parameters=124.4M, Context Length=10242025.12 | 3.247 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA) + SWANumber of Parameters=124.4M, Context Length=10242025.12 | 3.248 | — | — | — | — | — | — | — | — | — | — | |
| Transformer (LLaMA) + SWA + NSANumber of Parameters=125.4M, Context Length=40962025.12 | 3.248 | — | — | — | — | — | — | — | — | — | — | |
| mHCModel Scale=M2026.03 | 3.25 | — | — | — | — | — | — | — | — | — | — | |
| mHC-liteModel Scale=M2026.03 | 3.252 | — | — | — | — | — | — | — | — | — | — | |
| GatedFWANumber of Parameters=125.1M, Context Length=40962025.12 | 3.255 | — | — | — | — | — | — | — | — | — | — | |
| mHC-liteExperiment=Exp. 4 (single run)2026.05 | 3.2611 | — | — | — | — | — | — | — | — | 1.0615 | — | |
| HCModel Scale=M2026.03 | 3.264 | — | — | — | — | — | — | — | — | — | — | |
| mHC-liteExperiment=Exp. 3 (single run)2026.05 | 3.2656 | — | — | — | — | — | — | — | — | 1.063 | — | |
| msTBP-mHCExperiment=Exp. 3 (single run)2026.05 | 3.2691 | — | — | — | — | — | — | — | — | 1.0641 | — | |
| mHCExperiment=Exp. 4 (single run)2026.05 | 3.2695 | — | — | — | — | — | — | — | — | 1.0643 | — | |
| mHCExperiment=Exp. 3 (single run)2026.05 | 3.2714 | — | — | — | — | — | — | — | — | 1.0649 | — | |
| Transformer (LLaMA)Number of Parameters=124.4M, Context Length=40962025.12 | 3.273 | — | — | — | — | — | — | — | — | — | — | |
| lmaLTBP-mHCExperiment=Exp. 3 (single run)2026.05 | 3.2734 | — | — | — | — | — | — | — | — | 1.0655 | — | |
| Transformer (LLaMA) + SWANumber of Parameters=124.4M, Context Length=40962025.12 | 3.274 | — | — | — | — | — | — | — | — | — | — | |
| aTBP-mHCExperiment=Exp. 3 (single run)2026.05 | 3.2744 | — | — | — | — | — | — | — | — | 1.0659 | — | |
| KromHCExperiment=Exp. 3 (single run)2026.05 | 3.2759 | — | — | — | — | — | — | — | — | 1.0664 | — | |
| RWKVNumber of Parameters=124.4M, Context Length=40962025.12 | 3.276 | — | — | — | — | — | — | — | — | — | — | |
| oRTBP-mHCExperiment=Exp. 4 (single run)2026.05 | 3.276 | — | — | — | — | — | — | — | — | 1.0664 | — | |
| RTBP-mHCExperiment=Exp. 3 (single run)2026.05 | 3.2777 | — | — | — | — | — | — | — | — | 1.0667 | — | |
| TBP-mHCExperiment=Exp. 3 (single run)2026.05 | 3.2777 | — | — | — | — | — | — | — | — | 1.0669 | — | |
| sRTBP-mHCExperiment=Exp. 3 (single run)2026.05 | 3.2778 | — | — | — | — | — | — | — | — | 1.067 | — | |
| CRATE-α-smallModel size=57M2024.05 | 3.28 | — | — | — | — | — | — | — | — | — | — | |
| asTBP-mHCExperiment=Exp. 3 (single run)2026.05 | 3.2818 | — | — | — | — | — | — | — | — | 1.0683 | — | |
| Baseline GPT-2Params=103.0M, Step=22,0962026.04 | 3.2903 | 26.85 | — | — | — | — | — | — | — | — | — | |
| RWKVNumber of Parameters=124.4M, Context Length=10242025.12 | 3.291 | — | — | — | — | — | — | — | — | — | — | |
| KromHCExperiment=Exp. 4 (single run)2026.05 | 3.2965 | — | — | — | — | — | — | — | — | 1.073 | — | |
| RCModel Scale=M2026.03 | 3.328 | — | — | — | — | — | — | — | — | — | — | |
| RetNetNumber of Parameters=373.2M, Context Length=10242025.12 | 3.362 | — | — | — | — | — | — | — | — | — | — | |
| GLANumber of Parameters=123.8M, Context Length=40962025.12 | 3.364 | — | — | — | — | — | — | — | — | — | — |