Machine Translation on WMT En-De '14
30.15BLEUPATT-EG
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PATT-EG# Param.=387M, Speed=25.79, Backbone=M-BART2021.06 | 30.15 | — | — | |
| Transformer + CBMI-adaptiveModel Scale=big2022.03 | 30.12 | — | — | |
| MAT2020.06 | 29.9 | — | — | |
| Transformer + Self-Paced LearningModel Scale=big2022.03 | 29.85 | — | — | |
| Evolved Transformer2020.06 | 29.8 | — | — | |
| OmniNetvariant=P2021.03 | 29.8 | — | — | |
| PATT-EG# Param.=255M, Speed=49.54, Backbone=M-BERT2021.06 | 29.77 | — | — | |
| Transformer + Anti-Focal LossModel Scale=big2022.03 | 29.72 | — | — | |
| DynamicConv2020.06 | 29.7 | — | — | |
| Transformer + BMI-adaptiveModel Scale=big2022.03 | 29.69 | — | — | |
| Transformer + Freq-ExponentialModel Scale=big2022.03 | 29.66 | — | — | |
| Transformer + Focal LossModel Scale=big2022.03 | 29.65 | — | — | |
| Transformer + Freq-Chi-SquareModel Scale=big2022.03 | 29.64 | — | — | |
| 60L TransformerLayers=602021.03 | 29.5 | — | — | |
| TransformerModel Scale=big, Source=Re-implemented2022.03 | 29.31 | — | — | |
| Evolved Transformer2021.03 | 29.2 | — | — | |
| M-BART# Param.=610M, Speed=19.65, Backbone=M-BART2021.06 | 29.13 | — | — | |
| Transformer + CBMI-adaptiveModel Scale=base2022.03 | 29.01 | — | — | |
| Universal Transformersize=base, parameters=comparable to other base models2018.07 | 28.9 | — | — | |
| Weighted Transformer2020.06 | 28.9 | — | — | |
| Transformer + Self-Paced LearningModel Scale=base2022.03 | 28.69 | — | — | |
| Transformer + Anti-Focal LossModel Scale=base2022.03 | 28.65 | — | — | |
| Large TransformerSize=Large2021.03 | 28.6 | — | — | |
| Transformer + BMI-adaptiveModel Scale=base2022.03 | 28.56 | — | — | |
| Transformer + Freq-Chi-SquareModel Scale=base2022.03 | 28.47 | — | — | |
| Transformer + Freq-ExponentialModel Scale=base2022.03 | 28.43 | — | — | |
| Transformer + Focal LossModel Scale=base2022.03 | 28.43 | — | — | |
| Weighted Transformersize=base, parameters=comparable to other base models2018.07 | 28.4 | — | — | |
| TransformerModel Scale=big, Source=Original Paper2022.03 | 28.4 | — | — | |
| Transformer + LM PriorModel Scale=base2022.03 | 28.27 | — | — | |
| M-BERT# Param.=382M, Speed=26.51, Backbone=M-BERT2021.06 | 28.24 | — | — | |
| TransformerModel Scale=base, Source=Re-implemented2022.03 | 28.1 | — | — | |
| Transformersize=base, parameters=comparable to other base models2018.07 | 28 | — | — | |
| CODA#Params.=60.94M2021.05 | 28 | — | — | |
| Masked Label Smoothing (MLS)Vocabulary Sharing=true2022.03 | 27.91 | — | — | |
| Transformer + Simple FusionModel Scale=base2022.03 | 27.82 | — | — | |
| Knee scheduleBackbone=Transformers, Training Budget=70 epochs2020.03 | 27.53 | — | — | |
| TransformerVocabulary Sharing=false, Label Smoothing=true2022.03 | 27.53 | — | — | |
| TransformerVocabulary Sharing=true, Label Smoothing=false2022.03 | 27.51 | — | — | |
| TransformerVocabulary Sharing=true, Label Smoothing=true2022.03 | 27.44 | — | — | |
| Knee scheduleBackbone=Transformers, Training Budget=70 epochs, Explore strategy=Fixed 50%2020.03 | 27.41 | — | — | |
| TRANSFORMER#Params.=60.92M2021.05 | 27.4 | — | — | |
| Cosine DecayBackbone=Transformers, Training Budget=70 epochs2020.03 | 27.35 | — | — | |
| Transformer BASE#Params=65M2021.09 | 27.3 | — | — | |
| TransformerModel Scale=base, Source=Original Paper2022.03 | 27.3 | — | — | |
| BaselineBackbone=Transformers, Training Budget=70 epochs2020.03 | 27.29 | — | — | |
| Linear DecayBackbone=Transformers, Training Budget=70 epochs2020.03 | 27.29 | — | — | |
| TransformerVocabulary Sharing=false, Label Smoothing=false2022.03 | 27.21 | — | — | |
| TransEvolve-fullFF-2#Params=59M2021.09 | 27.2 | — | — | |
| One-Cycle DecayBackbone=Transformers, Training Budget=70 epochs2020.03 | 27.19 | — | — | |
| Universal Transformersize=small2018.07 | 26.8 | — | — | |
| 8-bit AdamW†Backbone=Transformer, Precision=8-bit, Quantization Policy=Do not quantize optimizer states for embedding layers2023.09 | 26.66 | — | — | |
| 32-bit AdamWBackbone=Transformer, Precision=32-bit2023.09 | 26.61 | — | — | |
| 32-bit AdafactorBackbone=Transformer, Precision=32-bit2023.09 | 26.52 | — | — | |
| 32-bit Adafactor+Backbone=Transformer, Precision=32-bit, beta1=02023.09 | 26.45 | — | — | |
| 4-bit FactorBackbone=Transformer, Precision=4-bit2023.09 | 26.45 | — | — | |
| 4-bit AdamWBackbone=Transformer, Precision=4-bit2023.09 | 26.28 | — | — | |
| Random Init# Param.=382M, Speed=25.642021.06 | 26.08 | — | — | |
| TransEvolve-fullFF-1#Params=53M2021.09 | 25.8 | — | — | |
| ARBeam width (b)=42018.02 | 24.57 | 44.9 | 7 | |
| DINOISERSampling=MBR-52022.12 | 24.3 | — | — | |
| TransEvolve-randomFF-2#Params=33M2021.09 | 23.8 | — | — | |
| ARBeam width (b)=12018.02 | 23.77 | 54 | 15.8 | |
| FlowSeq-largeKnowledge Distillation=true, Decoding=argmax2019.09 | 23.72 | — | — | |
| TransEvolve-randomFF-1#Params=27M2021.09 | 23.1 | — | — | |
| 32-bit SM3Backbone=Transformer, Precision=32-bit2023.09 | 22.72 | — | — | |
| LD4LG (MT5-base)Sampling=MBR-52022.12 | 22.4 | — | — | |
| Iterative RefinementRefinement steps (i_dec)=102018.02 | 21.61 | 90.4 | 12.3 | |
| Iterative RefinementAdaptive refinement steps=true2018.02 | 21.54 | 107.2 | 20.3 | |
| FlowSeq-baseKnowledge Distillation=true, Decoding=argmax2019.09 | 21.45 | — | — | |
| LD4LG (MT5-base)Sampling=Random2022.12 | 21.4 | — | — | |
| FlowSeq-largeKnowledge Distillation=false, Decoding=argmax2019.09 | 20.85 | — | — | |
| NAT-REGKnowledge Distillation=true, Decoding=argmax2019.09 | 20.65 | — | — | |
| Iterative RefinementRefinement steps (i_dec)=52018.02 | 20.26 | 139.7 | 23.1 | |
| CDCDSampling=MBR-102022.12 | 19.7 | — | — | |
| CDCDSampling=Random2022.12 | 19.3 | — | — | |
| NATFertility (FT)=true, NPD reranking=100 samples2018.02 | 19.17 | — | — | |
| FlowSeq-baseKnowledge Distillation=false, Decoding=argmax2019.09 | 18.55 | — | — | |
| CMLM-baseKnowledge Distillation=true, Decoding=argmax2019.09 | 18.12 | — | — | |
| NATFertility (FT)=true2018.02 | 17.69 | — | — | |
| NAT w/ FTKnowledge Distillation=true, Decoding=argmax2019.09 | 17.69 | — | — | |
| CTC LossKnowledge Distillation=true, Decoding=argmax2019.09 | 17.68 | — | — | |
| Iterative RefinementRefinement steps (i_dec)=22018.02 | 16.95 | 393.6 | 49.6 | |
| Diffusion-LMSampling=MBR-52022.12 | 15.3 | — | — | |
| CMLM-smallKnowledge Distillation=true, Decoding=argmax2019.09 | 15.06 | — | — | |
| Iterative RefinementRefinement steps (i_dec)=12018.02 | 13.91 | 511.4 | 83.3 | |
| NAT-IRKnowledge Distillation=true, Decoding=argmax2019.09 | 13.91 | — | — | |
| LV NARKnowledge Distillation=false, Decoding=argmax2019.09 | 11.8 | — | — | |
| CMLM-baseKnowledge Distillation=false, Decoding=argmax2019.09 | 10.88 | — | — |