Machine Translation on WMT newstest 2014
37.5Tokenized BLEUEnsemble of 8 LSTMs + PosUnk
Evaluation Results
| Method | Links | |
|---|---|---|
| Ensemble of 8 LSTMs + PosUnkVocab=80K, Corpus=36M2014.10 | 37.5 | |
| Jean et al. (2015) – 8 gated RNNs with search + UNK replacementVocab=500K, Corpus=12M2014.10 | 37.2 | |
| State of the art in WMT'14Vocab=All, Corpus=36M2014.10 | 37 | |
| Ensemble of 8 LSTMs + PosUnkVocab=40K, Corpus=12M2014.10 | 36.9 | |
| Sutskever et al. (2014) – 5 LSTMs, reranking 1000-best listsVocab=All, Corpus=12M2014.10 | 36.5 | |
| Ensemble of 8 LSTMsVocab=80K, Corpus=36M2014.10 | 35.6 | |
| Sutskever et al. (2014) – 5 LSTMsVocab=80K, Corpus=12M2014.10 | 34.8 | |
| Cho et al. (2014)– phrase table neural featuresVocab=All, Corpus=12M2014.10 | 34.5 | |
| Ensemble of 8 LSTMsVocab=40K, Corpus=12M2014.10 | 34.1 | |
| CYCLE (REV)#Params=343M, Training Data=Synthetic2021.04 | 33.54 | |
| Schwenk (2014) – neural language modelVocab=All, Corpus=12M2014.10 | 33.3 | |
| Single LSTM with 6 layers + PosUnkVocab=80K, Corpus=36M2014.10 | 33.1 | |
| Kiyono et al. (2020)#Params=514M, Training Data=Synthetic2021.04 | 33.1 | |
| Single LSTM with 6 layers + PosUnkVocab=40K, Corpus=12M2014.10 | 32.7 | |
| CYCLE#Params=242M, Training Data=Genuine2021.04 | 32.1 | |
| CYCLE (REV)#Params=242M, Training Data=Genuine2021.04 | 32.06 | |
| SEQUENCE#Params=242M, Training Data=Genuine2021.04 | 31.9 | |
| Single LSTM with 4 layers + PosUnkVocab=40K, Corpus=12M2014.10 | 31.8 | |
| Universal#Params=249M, Training Data=Genuine2021.04 | 31.73 | |
| Single LSTM with 6 layersVocab=80K, Corpus=36M2014.10 | 31.5 | |
| Vanilla#Params=242M, Training Data=Genuine2021.04 | 31.4 | |
| Single LSTM with 6 layersVocab=40K, Corpus=12M2014.10 | 30.4 | |
| mRASPSize=4.5M, Fine-tuning=true, Tokenization Protocol=tokenized2020.10 | 30.3 | |
| CTNMTSize=4.5M, Fine-tuning=true, Tokenization Protocol=tokenized, Setting=Transformer-base2020.10 | 30.1 | |
| Single LSTM with 4 layersVocab=40K, Corpus=12M2014.10 | 29.5 | |
| DirectSize=4.5M, Fine-tuning=true, Tokenization Protocol=tokenized2020.10 | 29.3 | |
| MASSSize=4.5M, Fine-tuning=true, Tokenization Protocol=tokenized2020.10 | 28.9 | |
| XLMSize=4.5M, Fine-tuning=true, Tokenization Protocol=tokenized2020.10 | 28.8 | |
| mBERTSize=4.5M, Fine-tuning=true, Tokenization Protocol=tokenized2020.10 | 28.6 | |
| Bahdanau et al. (2015) – single gated RNN with searchVocab=30K, Corpus=12M2014.10 | 28.5 |