Machine Translation on WMT English-German 2014 (test)
31.4BLEUKERMIT Joint → Unidirectional Finetuning
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| KERMIT Joint → Unidirectional FinetuningBidirectional capability=false, Iterations=≈ log2 n << 102019.06 | 31.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (Our Implementation)Bidirectional capability=false, Iterations=n2019.06 | 31.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Unidirectional (p(y | x) or p(x | y))Bidirectional capability=false, Iterations=≈ log2 n << 102019.06 | 30.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Data DiversificationBase Model=Scale Transformer [18]2019.11 | 30.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TransformerAugmentation=Multi-Agent [31]2019.11 | 30 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Dynamic Conv2019.11 | 29.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoE-TransformerMAC=62.5M, Parameters=138.2M2021.07 | 29.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (big)MAC=213.0M, Parameters=213.0M2021.07 | 29.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Scale Transformer2019.11 | 29.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Trans+Rel. Pos2019.11 | 29.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformervariant=Big, embedding_method=FRAGE2018.09 | 29.11 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer Base & Cutoff (w/ JS loss)Backbone=6-layer Transformer Base, JS loss=true, Cutoff strategy=token cutoff, Evaluation setting=case-sensitive tokenization and compound splitting2020.09 | 29.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Weighted TransformerModel size=large2017.11 | 28.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer Base & Cutoff (w/o JS loss)Backbone=6-layer Transformer Base, JS loss=false, Cutoff strategy=token cutoff, Evaluation setting=case-sensitive tokenization and compound splitting2020.09 | 28.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TransformerAugmentation=Ens-Distill [7]2019.11 | 28.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RNMT+architecture=multicol2019.05 | 28.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Joint → Unidirectional FinetuningBidirectional capability=false, Iterations=≈ log2 n << 102019.06 | 28.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RNMT+architecture=cascaded2019.05 | 28.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Joint + Marginal Refining (p(x) and p(y))Bidirectional capability=true, Iterations=≈ log2 n << 102019.06 | 28.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Joint → Bidirectional FinetuningBidirectional capability=true, Iterations=≈ log2 n << 102019.06 | 28.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DynamicConvbatch size=1, beam size=1, Param=200M, decoder layers=6, k=3,7,15,31,31,312019.01 | 28.5 | — | — | — | — | — | — | 110.9 | — | 3.9 | — | — | — | |
| RNMT+architecture=standard2019.05 | 28.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Weighted TransformerModel size=small2017.11 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TransformerModel size=large2017.11 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformersize=big2019.05 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformervariant=Big2018.09 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Evolved TransformerBackbone=6-layer Transformer Base2020.09 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Weighted TransformerBackbone=6-layer Transformer Base2020.09 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Adversarial TrainingBackbone=6-layer Transformer Base2020.09 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (big)Training Cost (FLOPs)=2.3 * 10^192017.06 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer2019.11 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TransformerAugmentation=Distill (T=S) [14]2019.11 | 28.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformervariant=Base, embedding_method=FRAGE2018.09 | 28.36 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (+SRU)number of layers=5, Model Size=90m2017.09 | 28.3 | — | — | — | — | — | — | 19 | 2.1 | — | — | — | — | |
| Transformer Base (So et al., 2019)Backbone=6-layer Transformer Base, Evaluation setting=case-sensitive tokenization and compound splitting, Source=Reported in (So et al., 2019)2020.09 | 28.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer-DRILLsize=base2019.05 | 28.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Joint → Bidirectional FinetuningBidirectional capability=true, Iterations=≈ log2 n << 102019.06 | 28.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (base)MAC=62.4M, Parameters=62.4M2021.07 | 28.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| AdminBackbone=6-layer Transformer Base2020.09 | 27.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (+SRU)number of layers=4, Model Size=79m2017.09 | 27.8 | — | — | — | — | — | — | 22 | 1.8 | — | — | — | — | |
| Transformer (Our Implementation)Bidirectional capability=false, Iterations=n2019.06 | 27.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Unidirectional (p(y | x) or p(x | y))Bidirectional capability=false, Iterations=≈ log2 n << 102019.06 | 27.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DynamicConvbatch size=1, beam size=1, Param=153M, decoder layers=3, k=3,7,152019.01 | 27.7 | — | — | — | — | — | — | 202.3 | — | 7.2 | — | — | — | |
| Transformer (base)number of layers=6, Model Size=76m2017.09 | 27.6 | — | — | — | — | — | — | 20 | 2 | — | — | — | — | |
| KERMIT Bidirectional (p(y | x) and p(x | y))Bidirectional capability=true, Iterations=≈ log2 n << 102019.06 | 27.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TransformerAugmentation=Distill (T>S) [14]2019.11 | 27.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PostNorm + LayerNormsource=this work2019.10 | 27.58 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PostNorm + FixNorm + ScaleNorm2019.10 | 27.57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer-Dualsize=base2019.05 | 27.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (Guo et al.)batch size=1, beam size=12019.01 | 27.4 | — | — | — | — | — | — | — | — | 1.6 | — | — | — | |
| Blockwise Parallel (Stern et al., 2018)Bidirectional capability=false, Iterations=≈ n/52019.06 | 27.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Insertion Transformer (Stern et al., 2019)Bidirectional capability=false, Iterations=≈ log2 n << 102019.06 | 27.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Joint (p(x, y))Bidirectional capability=true, Iterations=≈ log2 n << 102019.06 | 27.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TransformerModel size=small2017.11 | 27.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (Li et al.)batch size=1, beam size=12019.01 | 27.3 | — | — | — | — | — | — | — | — | 1.3 | — | — | — | |
| Transformersize=base2019.05 | 27.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (Vaswani et al., 2017)Bidirectional capability=false, Iterations=n2019.06 | 27.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PostNorm + LayerNormsource=published2019.10 | 27.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformervariant=Base2018.09 | 27.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer BaseBackbone=6-layer Transformer Base, Beam Decoding=Vaswani et al., 2017 configuration2020.09 | 27.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Transformer (base model)Training Cost (FLOPs)=3.3 * 10^182017.06 | 27.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Bidirectional (p(y | x) and p(x | y))Bidirectional capability=true, Iterations=≈ log2 n << 102019.06 | 27.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PreNorm + FixNorm + ScaleNorm2019.10 | 27.07 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PreNorm + LayerNorm2019.10 | 26.82 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvS2S EnsembleTraining Cost (FLOPs)=7.7 * 10^192017.06 | 26.36 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GNMT + RLensemble=true2019.05 | 26.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvS2Sensemble=true2019.05 | 26.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GNMT + RL EnsembleTraining Cost (FLOPs)=1.8 * 10^202017.06 | 26.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SliceNet-Super2018.06 | 26.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DynamicConvbatch size=1, beam size=1, Param=124M, decoder layers=1, k=312019.01 | 26.1 | — | — | — | — | — | — | 423 | — | 15.2 | — | — | — | |
| MoETraining Cost (FLOPs)=2.0 * 10^192017.06 | 26.03 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MoE2017.11 | 26 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MoE2019.05 | 26 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Joint + Marginal Refining (p(x) and p(y))Bidirectional capability=true, Iterations=≈ log2 n << 102019.06 | 25.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KERMIT Joint (p(x, y))Bidirectional capability=true, Iterations=≈ log2 n << 102019.06 | 25.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DenseNMT-En-De-152018.06 | 25.52 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SliceNet-Full2018.06 | 25.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Iterative Refinement (Lee et al., 2018)Bidirectional capability=false, Iterations=102019.06 | 25.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvS2S2017.11 | 25.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NART w/ hintsbatch size=1, beam size=1, B=4, candidates=92019.01 | 25.2 | — | — | — | — | — | — | — | — | 22.7 | — | — | — | |
| ConvS2S2018.06 | 25.16 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvS2S2018.09 | 25.16 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvS2STraining Cost (FLOPs)=9.6 * 10^182017.06 | 25.16 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ConvS2Sensemble=false2019.05 | 25.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Lookahead (Adam)Inner optimizer=ADAM, Training steps=50k, Model architecture=Transformer Base, Learning rate warm-up=True2019.07 | 24.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GNMT2018.06 | 24.61 | — | — | — | — | — | — | — | — | — | — | — | — | |
| WPM-32KCPU decoding time per sentence (s)=0.18822016.09 | 24.61 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GNMT+RL2017.11 | 24.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GNMT + RLensemble=false2019.05 | 24.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ADAMTraining steps=50k, Model architecture=Transformer Base, Learning rate warm-up=True2019.07 | 24.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GNMT + RLTraining Cost (FLOPs)=2.3 * 10^192017.06 | 24.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| AdafactorTraining steps=50k, Model architecture=Transformer Base, Learning rate warm-up=True2019.07 | 24.51 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Lookahead (Adam without warm-up)Inner optimizer=ADAM, Training steps=50k, Model architecture=Transformer Base, Learning rate warm-up=False2019.07 | 24.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| WPM-16KCPU decoding time per sentence (s)=0.19312016.09 | 24.36 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ENAT Embedding Mappingbatch size=1, beam size=1, rescoring candidates=92019.01 | 24.3 | — | — | — | — | — | — | — | — | 20.4 | — | — | — | |
| Mixed Word/CharacterCPU decoding time per sentence (s)=0.32682016.09 | 24.17 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Autoregressive (Lee et al.)batch size=1, beam size=12019.01 | 23.8 | — | — | — | — | — | — | 54 | — | — | — | — | — | |
| ByteNet2018.09 | 23.75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ByteNet2017.06 | 23.75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ByteNet2017.11 | 23.7 | — | — | — | — | — | — | — | — | — | — | — | — |