Machine Translation on IWSLT De-En 2014 (test)
38.61BLEUBiBERT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BiBERTbidirectional pretraining=true2022.06 | 38.61 | — | |
| Bi-SimCutbidirectional pretraining=true2022.06 | 38.37 | — | |
| CutoffLR Schedule=Knee, Training Epochs=1002020.03 | 37.78 | — | |
| CutoffLR Schedule=Knee (short budget), Training Epochs=702020.03 | 37.66 | — | |
| CutoffLR Schedule=Inv. Sqrt, Training Epochs=1002020.03 | 37.6 | — | |
| CipherDAug-2 keys|Θ|=1.2x2022.04 | 37.53 | — | |
| CipherDAug - 2 keyssrc aug=false, tgt aug=false, |Θ|=3x2022.04 | 37.53 | — | |
| BiBERT|Θ|=1x(+BERT)2022.04 | 37.5 | — | |
| CutoffLR Schedule=Inv. Sqrt (short budget), Training Epochs=702020.03 | 37.31 | — | |
| R-Drop2022.06 | 37.3 | — | |
| R-DROP|Θ|=1x2022.04 | 37.25 | — | |
| Data Diversesrc aug=true, tgt aug=true, |Θ|=7x2022.04 | 37 | — | |
| UniDropBackbone=Transformer2021.04 | 36.88 | — | |
| UniDrop|Θ|=1x2022.04 | 36.88 | — | |
| UniDrop2022.06 | 36.88 | — | |
| Document-level + BERTContext level=document-level, BERT-fused=true2020.02 | 36.69 | — | |
| MixReps+co-teaching2021.04 | 36.41 | — | |
| Mixed Rep2022.06 | 36.41 | — | |
| Mixed-Repr.src aug=true, tgt aug=true, |Θ|=2x2022.04 | 36.31 | — | |
| MUSEModel size=base2019.11 | 36.3 | — | |
| Mask Attention NetworksModel size=small, Parameters=37M2021.03 | 36.3 | — | |
| SRU++Param=20.4M, Training Time (Hrs)=8.5, Beam size=52021.02 | 36.3 | — | |
| MAT2021.04 | 36.22 | — | |
| MAT|Θ|=0.9x2022.04 | 36.22 | — | |
| RAML+SwitchOutsrc aug=true, tgt aug=true, |Θ|=1x2022.04 | 36.2 | — | |
| CipherDAug - 1 keysrc aug=false, tgt aug=false, |Θ|=2x2022.04 | 36.19 | — | |
| RAML+WordDropoutsrc aug=true, tgt aug=true, |Θ|=1x2022.04 | 36.13 | — | |
| Sentence-level + BERTContext level=sentence-level, BERT-fused=true2020.02 | 36.11 | — | |
| BERT-fused NMT2021.04 | 36.11 | — | |
| BERT Fuse|Θ|=1x(+BERT)2022.04 | 36.11 | — | |
| SRU++Param=19.6M, Training Time (Hrs)=7.5, Beam size=5, k=22021.02 | 36.1 | — | |
| RAMLsrc aug=false, tgt aug=true, |Θ|=1x2022.04 | 35.99 | — | |
| SwitchOutsrc aug=true, tgt aug=false, |Θ|=1x2022.04 | 35.9 | — | |
| TransformerParam=20.1M, Training Time (Hrs)=10.5, Beam size=52021.02 | 35.9 | — | |
| MUSE-simpleModel size=base2019.11 | 35.8 | — | |
| Soft Contextual Data Aug2021.04 | 35.78 | — | |
| Local Joint Self-attention2019.05 | 35.7 | — | |
| Local Joint Self-attention2019.11 | 35.7 | — | |
| IOT2021.04 | 35.62 | — | |
| WordDropoutsrc aug=true, tgt aug=false, |Θ|=1x2022.04 | 35.6 | — | |
| CONTBackbone=Transformer-small, Loss Function=N-Pairs loss, target-source representation similarity=true2022.05 | 35.55 | — | |
| Knee scheduleLR Schedule=Knee schedule, Architecture=Transformer, Training Epochs=502020.03 | 35.53 | — | |
| VAT2022.06 | 35.52 | — | |
| Macaron2019.11 | 35.4 | — | |
| Macaron2021.04 | 35.4 | — | |
| Macaron Net|Θ|=1x2022.04 | 35.4 | — | |
| Joint Self-attention2019.05 | 35.3 | — | |
| Cosine DecayLR Schedule=Cosine Decay, Architecture=Transformer, Training Epochs=502020.03 | 35.21 | — | |
| Wu et al. (2019)2019.05 | 35.2 | — | |
| DynamicConv2019.11 | 35.2 | — | |
| Dynamic ConvModel size=small2021.03 | 35.2 | — | |
| DynamicConv2021.04 | 35.2 | — | |
| Transformer w/ LABObeam size=5, Backbone=6-layer encoder-decoder transformer2023.05 | 35.2 | — | |
| Adversarial MLE2021.04 | 35.18 | — | |
| He et al. (2018)2019.05 | 35.1 | — | |
| BPE-Dropoutsrc aug=true, tgt aug=true, |Θ|=1x2022.04 | 35.1 | — | |
| Knee scheduleLR Schedule=Knee schedule, Training Budget=35 epochs, Checkpoint Selection=Best validation perplexity2020.03 | 35.08 | — | |
| Adam+ASAMArchitecture=Transformer, Base Optimizer=Adam2021.02 | 35.02 | — | |
| Transformer2022.06 | 34.99 | — | |
| Linear DecayLR Schedule=Linear Decay, Architecture=Transformer, Training Epochs=502020.03 | 34.97 | — | |
| Our Document-levelContext level=document-level, BERT-fused=false2020.02 | 34.95 | — | |
| AdamArchitecture=Transformer, Base Optimizer=Adam2021.02 | 34.86 | — | |
| Transformer2021.04 | 34.84 | — | |
| Adam+SAMArchitecture=Transformer, Base Optimizer=Adam2021.02 | 34.78 | — | |
| One-CycleLR Schedule=One-Cycle, Architecture=Transformer, Training Epochs=502020.03 | 34.77 | — | |
| Transformer|Θ|=44M2022.04 | 34.71 | — | |
| Standard enc-decComplexity (encoder/decoder)=O(n²)/O(n²), Encoder Attention=softmax, Decoder Attention=softmax2021.06 | 34.7 | — | |
| CONTBackbone=Transformer-small, Loss Function=N-Pairs loss, target-source representation similarity=false2022.05 | 34.69 | — | |
| Sentence-level baselineContext level=sentence-level, BERT-fused=false2020.02 | 34.64 | — | |
| Transformersrc aug=false, tgt aug=false, |Θ|=1x2022.04 | 34.64 | — | |
| Standard enc + PRF decComplexity (encoder/decoder)=O(n²)/O(n), Encoder Attention=softmax, Decoder Attention=PRF2021.06 | 34.6 | — | |
| Fixup2019.11 | 34.5 | — | |
| TransformerHardware-Aware=false, Hetero. Layers=false, Latency=3.3s, #Params=32M, FLOPS (G)=1.5, GPU Hours=2, CO2e (lbs)=5, Cloud Comp. Cost=$12 - $402020.05 | 34.5 | — | |
| HATHardware-Aware=true, Hetero. Layers=true, Latency=2.1s, #Params=23M, FLOPS (G)=1.1, GPU Hours=4, CO2e (lbs)=9, Cloud Comp. Cost=$24 - $802020.05 | 34.5 | — | |
| Transformer w/ LSbeam size=5, Backbone=6-layer encoder-decoder transformer2023.05 | 34.5 | — | |
| Naive CLBackbone=Transformer-small, Loss Function=N-Pairs loss, target-source representation similarity=true2022.05 | 34.47 | — | |
| Cosine DecayLR Schedule=Cosine Decay, Training Budget=35 epochs, Checkpoint Selection=Best validation perplexity2020.03 | 34.46 | — | |
| CONTBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 34.46 | — | |
| Naive CLBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 34.45 | — | |
| Transformer**Tokenization=BPE-based, Version=V2, Implementation=FairSeq, Number of parameters=52M2018.08 | 34.44 | — | |
| One-CycleLR Schedule=One-Cycle, Training Budget=35 epochs, Checkpoint Selection=Best validation perplexity2020.03 | 34.43 | — | |
| Transformer**Tokenization=BPE-based, Version=V1, Implementation=FairSeq, Number of parameters=46M2018.08 | 34.42 | — | |
| Dropout CLBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 34.41 | — | |
| Vaswani et al. (2017)2019.05 | 34.4 | — | |
| TransformerModel size=small, Parameters=36M2021.03 | 34.4 | — | |
| SSMBA CLBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 34.32 | — | |
| NPRF enc-dec w/ RPEComplexity (encoder/decoder)=O(nlogn)/O(nlogn), Relative Positional Encoding=true2021.06 | 34.3 | — | |
| TransformerKnowledge Distillation=true2021.03 | 34.29 | — | |
| Naive CLBackbone=Transformer-small, Loss Function=N-Pairs loss, target-source representation similarity=false2022.05 | 34.26 | — | |
| Transformer w/ CPbeam size=5, Backbone=6-layer encoder-decoder transformer2023.05 | 34.2 | — | |
| Transformer w/ AFLbeam size=5, Backbone=6-layer encoder-decoder transformer2023.05 | 34.2 | — | |
| Pervasive AttentionTokenization=BPE-based, Version=V2, Number of parameters=22M2018.08 | 34.18 | — | |
| ATIterations=N, latency (ms)=4862020.11 | 34.18 | — | |
| MLEBackbone=Transformer-small, Loss Function=MLE2022.05 | 34.18 | — | |
| Linear DecayLR Schedule=Linear Decay, Training Budget=35 epochs, Checkpoint Selection=Best validation perplexity2020.03 | 34.16 | — | |
| LATIterations=4, latency (ms)=732020.11 | 34.08 | — | |
| Miculicich et al. (2018)Context level=document-level, BERT-fused=false2020.02 | 33.97 | — | |
| Transformerbeam size=5, Backbone=6-layer encoder-decoder transformer2023.05 | 33.9 | — | |
| Pervasive AttentionTokenization=BPE-based, Version=V1, Number of parameters=11M2018.08 | 33.86 | — | |
| CeMATInitialization=CeMAT, Architecture=Mask-Predict, Decoding Strategy=Non-autoregressive2022.03 | 33.7 | — |