Language Modeling on PTB (val)
20.26PerplexityMDM-Prime-v2
Evaluation Results
| Method | Links | |
|---|---|---|
| MDM-Prime-v2zero-shot=true, ℓ=16, variant=*2026.03 | 20.26 | |
| MDM-Prime-v2zero-shot=true, ℓ=162026.03 | 24.5 | |
| BERT-Large-CAS2019.04 | 36.14 | |
| BERT-CAS2019.04 | 39.97 | |
| AWD-LSTM-MoSVocabulary=BERTVocab2019.04 | 43.47 | |
| MDM-Primezero-shot=true, ℓ=6, variant=*2026.03 | 44.49 | |
| GPT-CAS2019.04 | 46.24 | |
| AWD-LSTM-DOC (fin) x 5#Param=114M, fine-tuning=true, ensemble=true, ensemble_count=52018.08 | 48.63 | |
| AWD-LSTM-DOC x 5#Param=114M, ensemble=true, ensemble_count=52018.08 | 49.99 | |
| AWD-LSTM-MoSVocabulary=GPTVocab2019.04 | 50.2 | |
| Tensor-Transformercores=1, normalization=PowerNorm2020.03 | 51.6 | |
| MDM-Primezero-shot=true, ℓ=82026.03 | 53.77 | |
| MDM-Primezero-shot=true, ℓ=42026.03 | 53.98 | |
| AWD-LSTM-DOC (fin)#Param=23M, fine-tuning=true2018.08 | 54.12 | |
| Tensor-Transformercores=22020.03 | 54.3 | |
| AWD-LSTM-DOC#Param=23M2018.08 | 54.62 | |
| Tensor-Transformercores=12020.03 | 55.4 | |
| AWD-LSTM-MoS#Param=22M2018.08 | 56.54 | |
| Transformer-XLsize=base2020.03 | 56.7 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=0.5, Dropout rate multiplier (λ)=x0.82018.05 | 57.1 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=1, Dropout rate multiplier (λ)=x0.82018.05 | 57.1 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=1, Dropout rate multiplier (λ)=x0.82018.05 | 57.3 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=0.5, Dropout rate multiplier (λ)=x0.92018.05 | 57.3 | |
| DETSoftmax Temperature (Temp)=opt2018.05 | 57.5 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=1, Dropout rate multiplier (λ)=x0.92018.05 | 57.5 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=0, Dropout rate multiplier (λ)=x0.82018.05 | 57.5 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=1, Dropout rate multiplier (λ)=x0.92018.05 | 57.5 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=0.5, Dropout rate multiplier (λ)=x1.02018.05 | 57.8 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=0.5, Dropout rate multiplier (λ)=x0.92018.05 | 57.9 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=0, Dropout rate multiplier (λ)=x0.92018.05 | 57.9 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=0.5, Dropout rate multiplier (λ)=x1.02018.05 | 58 | |
| Tensor-Transformercores=1, normalization=LayerNorm, re-implemented=true2020.03 | 58 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=0.5, Dropout rate multiplier (λ)=x0.82018.05 | 58.1 | |
| LSTM + 15 softmax expertsParams (M)=22, Search Method=manual2018.06 | 58.1 | |
| DARTS (second order)Params (M)=23, Search Cost (GPU days)=1, #ops=4, Search Method=gradient-based2018.06 | 58.1 | |
| AWD-LSTM-MoS2020.03 | 58.1 | |
| DARTS 2ndEvaluation Epochs=8000, Parameters (M)=23, Search Time (GPU Days)=12021.06 | 58.1 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=0, Dropout rate multiplier (λ)=x1.02018.05 | 58.3 | |
| AMCSoftmax Temperature (Temp)=opt, Alpha (α)=1, Dropout rate multiplier (λ)=x1.02018.05 | 58.4 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=1, Dropout rate multiplier (λ)=x1.02018.05 | 58.5 | |
| AWD-LSTM + Fraternal Dropout#Param=24M2018.08 | 58.9 | |
| Adaptive Input2020.03 | 59.1 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=0, Dropout rate multiplier (λ)=x0.82018.05 | 59.6 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=0, Dropout rate multiplier (λ)=x0.92018.05 | 59.7 | |
| AMCSoftmax Temperature (Temp)=1, Alpha (α)=0, Dropout rate multiplier (λ)=x1.02018.05 | 59.7 | |
| Tensor-Transformercores=1, normalization=PowerNorm-V2020.03 | 59.7 | |
| GDASEvaluation Epochs=2000, Parameters (M)=23, Search Time (GPU Days)=0.42021.06 | 59.8 | |
| AWD-LSTM#Param=24M2018.08 | 60 | |
| AWD-LSTM2022.01 | 60 | |
| DARTS (first order)Params (M)=23, Search Cost (GPU days)=0.5, #ops=4, Search Method=gradient-based2018.06 | 60.2 | |
| DARTS 1stEvaluation Epochs=8000, Parameters (M)=23, Search Time (GPU Days)=0.52021.06 | 60.2 | |
| LSTMParams (M)=24, Search Method=manual2018.06 | 60.7 | |
| ENAS (Pham et al., 2018b)+Params (M)=24, Search Cost (GPU days)=0.5, #ops=4, Search Method=RL2018.06 | 60.8 | |
| ENAS++Evaluation Epochs=8000, Parameters (M)=24, Search Time (GPU Days)=0.5, Note=Reevaluated by Liu et al. 20192021.06 | 60.8 | |
| LSTM with skip connections#Param=24M2018.08 | 60.9 | |
| DETSoftmax Temperature (Temp)=12018.05 | 60.9 | |
| LSTM + skip connectionsParams (M)=24, Search Method=manual2018.06 | 60.9 | |
| fp32Precision Format=FP32, Model Architecture=LSTM2018.04 | 61.31 | |
| HBFPMantissa Bits=12, Weight Storage Bits=16, Tile Size=24, Arithmetic Precision=12-bit, Model Architecture=LSTM2018.04 | 61.35 | |
| Random search baseline+Params (M)=23, Search Cost (GPU days)=2, #ops=4, Search Method=random2018.06 | 61.8 | |
| Random searchEvaluation Epochs=8000, Parameters (M)=23, Search Time (GPU Days)=22021.06 | 61.8 | |
| HBFPMantissa Bits=8, Weight Storage Bits=16, Tile Size=24, Arithmetic Precision=8-bit, Model Architecture=LSTM2018.04 | 61.86 | |
| MDM-Primezero-shot=true, ℓ=62026.03 | 62.62 | |
| Variational RHN + WT + IOG#Param=29M2018.08 | 67 | |
| Variational RHN + WT#Param=23M2018.08 | 67.9 | |
| Variational RHNParams (M)=23, Search Method=manual2018.06 | 67.9 | |
| ENAS (Pham et al., 2018b)*Params (M)=24, Search Cost (GPU days)=0.5, #ops=4, Search Method=RL2018.06 | 68.3 | |
| Variational RHN#Param=32M2018.08 | 71.2 | |
| Tensor-Transformercores=1, normalization=BatchNorm2020.03 | 71.7 | |
| BERT2019.04 | 72.99 | |
| MDM-Primezero-shot=true, ℓ=22026.03 | 74.81 | |
| Tied-LSTM2020.03 | 75.7 | |
| Variational LSTM (large)#Param=66M2018.08 | 77.9 | |
| GPT2019.04 | 79.44 | |
| ARParam=0.1B, Evaluation Protocol=Zero-shot2026.05 | 81.07 | |
| Variational LSTM (medium)#Param=20M2018.08 | 81.9 | |
| AR TransformerTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 82 | |
| ARMzero-shot=true2026.03 | 82.05 | |
| LSTM (large)#Param=66M2018.08 | 82.2 | |
| LSTM (medium)#Param=20M2018.08 | 86.2 | |
| MDMzero-shot=true, variant=*2026.03 | 87.29 | |
| DCDMParam=0.1B, Training Tokens=128B, Evaluation Protocol=Zero-shot2026.05 | 87.82 | |
| Duozero-shot=true2026.03 | 89.35 | |
| EDLM-coARzero-shot=true2026.03 | 89.73 | |
| LoopMDMS (inference-time loop count)=12, Zero-shot protocol=true, Model parameters=170M2026.05 | 90.2 | |
| MDLMParam=0.1B, Training Tokens=256B, Evaluation Protocol=Zero-shot2026.05 | 90.96 | |
| EDLM-NCEzero-shot=true2026.03 | 93.21 | |
| MDMzero-shot=true2026.03 | 95.26 | |
| BD3-LMTraining steps=250K, Training Dataset=OWT, Zero-shot=true, L'=162025.06 | 95.87 | |
| BD3-LMzero-shot=true2026.03 | 96.81 | |
| BDLMParam=0.1B, Training Tokens=256B, Evaluation Protocol=Zero-shot2026.05 | 96.81 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.1252025.06 | 97.46 | |
| SEDD AbsorbTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 99.59 | |
| SEDDzero-shot=true2026.03 | 100.09 | |
| MDLMTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 100.17 | |
| LoopMDMS (inference-time loop count)=6, Zero-shot protocol=true, Model parameters=170M2026.05 | 100.2 | |
| ARMzero-shot=true, variant=*2026.03 | 103.25 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.252025.06 | 105.19 | |
| MDMZero-shot protocol=true, Model parameters=170M2026.05 | 108.5 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.52025.06 | 110.7 |