Machine Translation on IWSLT De-En 2014
36.11BLEUBERT-fused model
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BERT-fused model2020.02 | 36.11 | — | |
| Multi-agent dual learning2020.02 | 35.56 | — | |
| Tied-Transformer2020.02 | 35.52 | — | |
| Adversarial TrainingBackbone=Transformer, Size=Base2019.06 | 35.18 | — | |
| Loss to teach2020.02 | 34.8 | — | |
| Role-interactive layer2020.02 | 34.74 | — | |
| TransformerSize=Base2019.06 | 34.43 | — | |
| Variational attention2020.02 | 33.68 | — | |
| Adversarial TrainingBackbone=Transformer, Size=Small2019.06 | 33.61 | — | |
| TransformerSize=Small2019.06 | 32.47 | — | |
| CNN-aArchitecture=CNN2019.06 | 30.04 | — | |
| Actor-criticArchitecture=Actor-critic2019.06 | 28.53 | — | |
| FlowSeq-baseKnowledge Distillation=true, Decoding=argmax2019.09 | 27.55 | — | |
| FlowSeq-baseKnowledge Distillation=false, Decoding=argmax2019.09 | 24.75 | — | |
| NAT-IRKnowledge Distillation=true, Decoding=argmax2019.09 | 21.86 | — | |
| NAT w/ FTKnowledge Distillation=true, Decoding=argmax2019.09 | 20.32 | — | |
| Deng et al.2019.10 | — | 33.08 | |
| LightConvParam=38.14M2019.10 | — | 34.84 | |
| LightConv + CGC EncoderParam=38.15M2019.10 | — | 35.21 | |
| LightConv + Dynamic EncoderParam=38.44M2019.10 | — | 35.03 | |
| TransformerParam=39.47M2019.10 | — | 34.41 |