Unconditional Text Generation on OpenWebText
1.07Gen. PPLGreedy
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GreedyCategory=AR, NFE=10242026.06 | 1.07 | — | 1.65 | — | |
| ARSampling steps (T)=32, Nucleus sampling=true2025.09 | 1.21 | 0.76 | 5.22 | — | |
| ARSampling steps (T)=64, Nucleus sampling=true2025.09 | 1.21 | 0.76 | 5.22 | — | |
| ARSampling steps (T)=128, Nucleus sampling=true2025.09 | 1.21 | 0.76 | 5.22 | — | |
| ProSeCoT=10242026.02 | 11.1 | 0.523 | 5.06 | — | |
| IMDMSteps=64, Number of Parameters=860M, Entropy-matched=false2026.05 | 11.6 | — | 4.95 | — | |
| ART=10242026.02 | 12.1 | 0.76 | 5.22 | — | |
| TransformerSteps=1024, p=0.92026.02 | 12.11 | — | 5.24 | — | |
| ProSeCoT=5122026.02 | 12.5 | 0.449 | 5.12 | — | |
| IMDMSteps=32, Number of Parameters=860M, Entropy-matched=false2026.05 | 13.95 | — | 5.06 | — | |
| AR (student-size)Training Steps=1M, Sequence Length=1024, Reference Type=Autoregressive2026.03 | 14.1 | 0.691 | — | — | |
| Decoded noise2026.03 | 14.24 | — | — | 0.038 | |
| Data (ref.)Sampling steps (T)=322025.09 | 14.8 | 1 | 5.44 | — | |
| Data (ref.)Sampling steps (T)=642025.09 | 14.8 | 1 | 5.44 | — | |
| Data (ref.)Sampling steps (T)=1282025.09 | 14.8 | 1 | 5.44 | — | |
| ProSeCoT=2562026.02 | 14.9 | 0.419 | 5.19 | — | |
| IMDMSteps=64, Number of Parameters=860M, Entropy-matched=true2026.05 | 14.9 | — | 5.14 | — | |
| PRISMT=10242026.02 | 15.3 | 0.527 | 5.1 | — | |
| PRISMT=5122026.02 | 16.4 | 0.423 | 5.12 | — | |
| Training set2026.03 | 16.75 | — | — | 0.2191 | |
| IMDMSteps=32, Number of Parameters=860M, Entropy-matched=true2026.05 | 16.9 | — | 5.18 | — | |
| PRISMT=2562026.02 | 18 | 0.294 | 5.15 | — | |
| TransformerSteps=1024, p=0.952026.02 | 18.51 | — | 5.4 | — | |
| IMDMSteps=16, Number of Parameters=860M, Entropy-matched=false2026.05 | 19.07 | — | 5.18 | — | |
| ProSeCoT=1282026.02 | 19.7 | 0.167 | 5.26 | — | |
| IDLM-MDLMSteps=322026.02 | 20.37 | — | 5.23 | — | |
| ReMDMT=5122026.02 | 21.1 | 0.35 | 5.21 | — | |
| PRISMT=1282026.02 | 21.5 | 0.118 | 5.18 | — | |
| IMDMSteps=16, Number of Parameters=860M, Entropy-matched=true2026.05 | 21.66 | — | 5.24 | — | |
| IMDMSampling Steps=642026.05 | 23.31 | — | 5.18 | — | |
| Manta-LMSequence Length=2562026.05 | 23.56 | — | 5.9714 | — | |
| Manta-LMSequence Length=1282026.05 | 23.8 | — | 6.0382 | — | |
| PAPLSampling steps (T)=128, Nucleus sampling=true2025.09 | 24.33 | 0.067 | 5.16 | — | |
| BD3-LM + Gumbel DistillationTraining Steps=1M, Sequence Length=1024, L'=4, Teacher=GPT-2-Large2026.03 | 24.37 | 0.304 | — | — | |
| Recovered training set2026.03 | 25.07 | — | — | 0.2282 | |
| Di4CSampling Steps=642026.05 | 25.8 | — | 5.27 | — | |
| BD3-LMTraining Steps=1M, Sequence Length=1024, L'=42026.03 | 26.4 | 0.251 | — | — | |
| MCDLM + SDTTPretrain Steps=100k, Distill Steps=50k, Sampling steps=1024, Sampling precision=FP64, Sampler type=ancestral2026.04 | 27.1 | — | 5.2 | — | |
| IMDMSampling Steps=322026.05 | 27.41 | — | 5.24 | — | |
| SDTTSteps=64, Number of Parameters=860M, Entropy-matched=false2026.05 | 27.98 | — | 5.13 | — | |
| MCDLM + SDTTPretrain Steps=100k, Distill Steps=50k, Sampling steps=512, Sampling precision=FP64, Sampler type=ancestral2026.04 | 28.1 | — | 4.9 | — | |
| ReMDMT=10242026.02 | 28.6 | 0.403 | 5.38 | — | |
| Manta-LMSequence Length=642026.05 | 29.18 | — | 5.7981 | — | |
| PAPLSampling steps (T)=64, Nucleus sampling=true2025.09 | 29.98 | 0.046 | 5.24 | — | |
| ReMDMT=2562026.02 | 30.5 | 0.216 | 5.34 | — | |
| MDLM - SDTTPretrain Steps=100k, Distill Steps=50k, Sampling steps=1024, Sampling precision=FP64, Sampler type=ancestral2026.04 | 31.2 | — | 5.4 | — | |
| Di4CSampling Steps=322026.05 | 31.38 | — | 5.32 | — | |
| Manta-LMSequence Length=322026.05 | 32.1 | — | 6.0459 | — | |
| MCDLM–PPLOptimizedPretrain Steps=150k, Distill Steps=0, Sampling steps=256, Sampling precision=FP64, Sampler type=ancestral2026.04 | 32.5 | — | 5.3 | — | |
| RADDSequence Length=2562026.05 | 32.68 | — | 6.4214 | — | |
| IDLM-MDLMSteps=162026.02 | 32.74 | — | 5.42 | — | |
| SDTTSampling Steps=642026.05 | 33.21 | — | 5.3 | — | |
| MCDLM–PPLOptimizedPretrain Steps=150k, Distill Steps=0, Sampling steps=512, Sampling precision=FP64, Sampler type=ancestral2026.04 | 33.5 | — | 5.2 | — | |
| IMDMSteps=8, Number of Parameters=860M, Entropy-matched=false2026.05 | 33.55 | — | 5.28 | — | |
| MCDLM–PPLOptimizedPretrain Steps=150k, Distill Steps=0, Sampling steps=1024, Sampling precision=FP64, Sampler type=ancestral2026.04 | 33.8 | — | 5.4 | — | |
| MDLM + Gumbel DistillationTraining Steps=1M, Sequence Length=1024, Teacher=GPT-2-Large2026.03 | 34.33 | 0.282 | — | — | |
| TEncDM Enc + QS0.4 + SC0.5Model Type=Encoder, Self-conditioning=p=0.5, q-sampling=t_act/T=0.42026.04 | 34.4 | 0.732 | — | 0.204 | |
| SampleCategory=AR, NFE=10242026.06 | 35.45 | — | 5.58 | — | |
| IMDMSampling Steps=162026.05 | 35.48 | — | 5.3 | — | |
| SDTTSteps=32, Number of Parameters=860M, Entropy-matched=false2026.05 | 36.03 | — | 5.18 | — | |
| TransformerSteps=1024, p=1.02026.02 | 36.45 | — | 5.6 | — | |
| RADDSequence Length=1282026.05 | 37.53 | — | 6.4133 | — | |
| MDLM+DFMSampling steps (T)=128, Nucleus sampling=true2025.09 | 37.9 | 0.041 | 5.31 | — | |
| MCDLM–PPLOptimizedPretrain Steps=150k, Distill Steps=0, Sampling steps=128, Sampling precision=FP64, Sampler type=ancestral2026.04 | 38.1 | — | 5.2 | — | |
| MDLMTraining Steps=1M, Sequence Length=10242026.03 | 38.34 | 0.217 | — | — | |
| IDLM-DCDgSteps=322026.02 | 38.57 | — | 5.35 | — | |
| TEncDM Enc + SC0.5Model Type=Encoder, Self-conditioning=p=0.52026.04 | 38.6 | 0.716 | — | 0.217 | |
| MCDLM–PPLOptimizedPretrain Steps=150k, Distill Steps=0, Sampling steps=64, Sampling precision=FP64, Sampler type=ancestral2026.04 | 40.1 | — | 5.3 | — | |
| PAPLSampling steps (T)=32, Nucleus sampling=true2025.09 | 40.19 | 0.013 | 5.32 | — | |
| ARPretrain Steps=75K, Distill Steps=0, Sampling steps=1024, Sampling precision=FP64, Sampler type=ancestral2026.04 | 40.2 | — | 5.6 | — | |
| SDTTSampling Steps=322026.05 | 40.41 | — | 5.34 | — | |
| MDLMSteps=10242026.02 | 41.29 | — | 5.28 | — | |
| IDLM-DCDaSteps=322026.02 | 42.03 | — | 5.41 | — | |
| ReMDMT=1282026.02 | 42.5 | 0.057 | 5.43 | — | |
| ReMDMSampling steps (T)=128, Nucleus sampling=true2025.09 | 42.5 | 0.057 | 5.43 | — | |
| MDLM+FBSampling steps (T)=128, Nucleus sampling=true2025.09 | 42.8 | 0.064 | 5.44 | — | |
| IDLM-DCDgSteps=162026.02 | 43.21 | — | 5.41 | — | |
| SEDD AbsorbSteps=10242026.02 | 43.31 | — | 5.25 | — | |
| Di4CSampling Steps=162026.05 | 44.12 | — | 5.37 | — | |
| RADDSequence Length=642026.05 | 46.23 | — | 6.3501 | — | |
| Duo-DCDgSteps=322026.02 | 46.31 | — | 5.38 | — | |
| Manta-LMSequence Length=162026.05 | 46.92 | — | 6.1232 | — | |
| TS-DFMSize (B)=0.17, Tokens (B)=78.7+6.4‡, Sampling Steps=16, Fine-tuning=true2026.05 | 47.2 | — | 7.4 | — | |
| CoDARTemperature (T)=0.00, Sampling steps=250, Number of samples=10002026.03 | 47.71 | — | — | 0.166 | |
| MCDLM–PPLOptimizedPretrain Steps=150k, Distill Steps=0, Sampling steps=32, Sampling precision=FP64, Sampler type=ancestral2026.04 | 48.7 | — | 5.3 | — | |
| CoDARTemperature (T)=0.25, Sampling steps=250, Number of samples=10002026.03 | 50.68 | — | — | 0.1937 | |
| MDLMT=10242026.02 | 51.3 | 0.042 | 5.46 | — | |
| IDLM-DCDaSteps=162026.02 | 51.86 | — | 5.44 | — | |
| MDLMT=5122026.02 | 53 | 0.031 | 5.48 | — | |
| MCDLMPretrain Steps=150k, Distill Steps=0, Sampling steps=1024, Sampling precision=FP64, Sampler type=ancestral2026.04 | 53.4 | — | 5.5 | — | |
| IDLM-DCDgSteps=82026.02 | 53.55 | — | 5.41 | — | |
| SDTT (7 rounds)Size (B)=0.86, Sampling Steps=16, Fine-tuning=false2026.05 | 53.6 | — | 7.7 | — | |
| DuoSize (B)=0.17, Tokens (B)=524, Sampling Steps=16, Fine-tuning=false2026.05 | 53.6 | — | 7.7 | — | |
| TEncDM Emb + QS0.4 + SC0.5Model Type=Embedding, Self-conditioning=p=0.5, q-sampling=t_act/T=0.42026.04 | 53.9 | 0.673 | — | 0.234 | |
| IDLM-DuogSteps=322026.02 | 54.05 | — | 5.49 | — | |
| Duo-DCDgSteps=162026.02 | 54.11 | — | 5.37 | — | |
| MCDLMPretrain Steps=150k, Distill Steps=0, Sampling steps=256, Sampling precision=FP64, Sampler type=ancestral2026.04 | 55.4 | — | 5.5 | — | |
| MDLMT=2562026.02 | 55.8 | 0.023 | 5.49 | — | |
| TS-DFMSize (B)=0.17, Tokens (B)=78.7+6.4‡, Sampling Steps=8, Fine-tuning=true2026.05 | 56.1 | — | 7.2 | — | |
| SDTTSteps=16, Number of Parameters=860M, Entropy-matched=false2026.05 | 57.28 | — | 5.24 | — |