Machine Translation on WMT En-Fr newstest 2014 (test)
43.4BLEURevCol-Transformer
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RevCol-TransformerEncoder arch=B=(1,1,1,1), COL=4, Encoder dmodel=768, Encoder dff=3072, Encoder head=12, Decoder arch=N=6, Decoder dmodel=768, Decoder dff=3072, Decoder head=12, Params=209M2022.12 | 43.4 | — | |
| Ott et al. (2018)Param (En-De)=210M2019.01 | 43.2 | — | |
| DynamicConvParam (En-De)=213M2019.01 | 43.2 | — | |
| Ott et al.2020.02 | 43.2 | — | |
| Wu et al.2020.02 | 43.2 | — | |
| TaLK Convolution2020.02 | 43.2 | — | |
| LightConvParam (En-De)=202M2019.01 | 43.1 | — | |
| Transformer bigEncoder arch=N=6, Encoder dmodel=1024, Encoder dff=4096, Encoder head=16, Decoder arch=N=6, Decoder dmodel=1024, Decoder dff=4096, Decoder head=16, Params=221M2022.12 | 43.07 | — | |
| Shaw et al. (2018)2019.01 | 41.5 | — | |
| Shaw et al.2020.02 | 41.5 | — | |
| Ahmed et al. (2017)Param (En-De)=213M2019.01 | 41.4 | — | |
| Ahmed et al.2020.02 | 41.4 | — | |
| Vaswani et al. (2017)Param (En-De)=213M2019.01 | 41 | — | |
| Chen et al. (2018)Param (En-De)=379M2019.01 | 41 | — | |
| Vaswani et al.2020.02 | 41 | — | |
| Chen et al.2020.02 | 41 | — | |
| MoE with 2048 Experts (longer training)ops/timestep=85M, Total #Parameters=8.7B, Training Time=6 days/64 k40s2017.01 | 40.56 | 2.63 | |
| Gehring et al. (2017)Param (En-De)=216M2019.01 | 40.5 | — | |
| Gehring et al.2020.02 | 40.5 | — | |
| MoE with 2048 Expertsops/timestep=85M, Total #Parameters=8.7B, Training Time=3 days/64 k40s2017.01 | 40.35 | 2.69 | |
| GNMT+RLops/timestep=214M, Total #Parameters=278M, Training Time=6 days/96 k80s2017.01 | 39.92 | 2.96 | |
| GNMTops/timestep=214M, Total #Parameters=278M, Training Time=6 days/96 k80s2017.01 | 39.22 | 2.79 | |
| DeepAtt+PosUnk2017.01 | 39.2 | — | |
| GNMTSupervision=Supervised2017.10 | 38.95 | — | |
| DeepAtt2017.01 | 37.7 | — | |
| PBMT2017.01 | 37 | — | |
| LSTM (6-layer+PosUnk)2017.01 | 33.1 | — | |
| LSTM (6-layer)2017.01 | 31.5 | — | |
| Proposed (full) + 100k parallelSupervision=Semi-supervised, Parallel data=100k2017.10 | 21.81 | — | |
| Proposed (full) + 100k parallelSupervision=Semi-supervised, Parallel data=100k2017.10 | 21.74 | — | |
| Comparable NMT (full parallel)Supervision=Supervised, Parallel data=full2017.10 | 20.48 | — | |
| Comparable NMT (full parallel)Supervision=Supervised, Parallel data=full2017.10 | 19.89 | — | |
| Proposed (full) + 10k parallelSupervision=Semi-supervised, Parallel data=10k2017.10 | 18.57 | — | |
| Proposed (full) + 10k parallelSupervision=Semi-supervised, Parallel data=10k2017.10 | 17.34 | — | |
| Proposed (+ backtranslation)Supervision=Unsupervised, Components=denoising, backtranslation2017.10 | 15.56 | — | |
| Proposed (+ BPE)Supervision=Unsupervised, Components=denoising, backtranslation, BPE2017.10 | 15.56 | — | |
| Proposed (+ backtranslation)Supervision=Unsupervised, Components=denoising, backtranslation2017.10 | 15.13 | — | |
| Proposed (+ BPE)Supervision=Unsupervised, Components=denoising, backtranslation, BPE2017.10 | 14.36 | — | |
| Comparable NMT (100k parallel)Supervision=Supervised, Parallel data=100k2017.10 | 10.4 | — | |
| Baseline (emb. nearest neighbor)Supervision=Unsupervised2017.10 | 9.98 | — | |
| Comparable NMT (100k parallel)Supervision=Supervised, Parallel data=100k2017.10 | 9.19 | — | |
| Proposed (denoising)Supervision=Unsupervised, Components=denoising2017.10 | 7.28 | — | |
| Baseline (emb. nearest neighbor)Supervision=Unsupervised2017.10 | 6.25 | — | |
| Proposed (denoising)Supervision=Unsupervised, Components=denoising2017.10 | 5.33 | — | |
| Comparable NMT (10k parallel)Supervision=Supervised, Parallel data=10k2017.10 | 1.88 | — | |
| Comparable NMT (10k parallel)Supervision=Supervised, Parallel data=10k2017.10 | 1.66 | — |