Machine Translation on WMT16 German-English (test)
40.6BLEUGPT-3
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT-3Few-shot learning=Few-shot2020.05 | 40.6 | — | — | — | — | |
| SOTA (Supervised)Training=Supervised2020.05 | 40.2 | — | — | — | — | |
| MASSTraining=Unsupervised2020.05 | 35.2 | — | — | — | — | |
| Supervised NMTUSMT Tuning=Supervised, Sentence pairs=5.6M2018.10 | 34.9 | — | — | — | — | |
| XLMEncoder Pre-training=MLM, Decoder Pre-training=MLM2019.01 | 34.3 | — | — | — | — | |
| XLMTraining=Unsupervised2020.05 | 34.3 | — | — | — | — | |
| mBARTTraining=Unsupervised2020.05 | 34 | — | — | — | — | |
| Supervised NMTUSMT Tuning=Supervised, Sentence pairs=2.8M2018.10 | 33.8 | — | — | — | — | |
| XLMEncoder Pre-training=MLM, Decoder Pre-training=-2019.01 | 33.2 | — | — | — | — | |
| XLMEncoder Pre-training=MLM, Decoder Pre-training=CLM2019.01 | 32.9 | — | — | — | — | |
| Supervised NMTUSMT Tuning=Supervised, Sentence pairs=1.4M2018.10 | 32.5 | — | — | — | — | |
| XLMEncoder Pre-training=CLM, Decoder Pre-training=MLM2019.01 | 32.5 | — | — | — | — | |
| XLMEncoder Pre-training=CLM, Decoder Pre-training=CLM2019.01 | 30.5 | — | — | — | — | |
| GPT-3Few-shot learning=One-shot2020.05 | 30.4 | — | — | — | — | |
| XLMEncoder Pre-training=-, Decoder Pre-training=CLM2019.01 | 30.3 | — | — | — | — | |
| UNMT (this work)USMT Tuning=Supervised, Filtering=Enabled2018.10 | 28.8 | — | — | — | — | |
| XLMEncoder Pre-training=MLM, Decoder Pre-training=-2019.01 | 28.6 | — | — | — | — | |
| UNMT (this work) w/o filteringUSMT Tuning=Supervised2018.10 | 28.2 | — | — | — | — | |
| XLMEncoder Pre-training=EMB, Decoder Pre-training=EMB2019.01 | 27.3 | — | — | — | — | |
| GPT-3Few-shot learning=Zero-shot2020.05 | 27.2 | — | — | — | — | |
| UNMT (this work) w/o filteringUSMT Tuning=No2018.10 | 27 | — | — | — | — | |
| Supervised trainingTraining mode=Supervised, Loss=Standard cross-entropy2018.04 | 26.99 | — | — | — | — | |
| UNMT (this work)USMT Tuning=No, Filtering=Enabled2018.10 | 26.7 | — | — | — | — | |
| XLMEncoder Pre-training=CLM, Decoder Pre-training=-2019.01 | 26 | — | — | — | — | |
| Lample et al. (2018b) USMT+UNMTUSMT Tuning=No2018.10 | 25.2 | — | — | — | — | |
| PBSMT + NMT2019.01 | 25.2 | — | — | — | — | |
| Artetxe et al. (2018b) USMTUSMT Tuning=back-translation2018.10 | 23.1 | — | — | — | — | |
| Lample et al. (2018b) USMTUSMT Tuning=No2018.10 | 22.7 | — | — | — | — | |
| PBSMT2019.01 | 22.7 | — | — | — | — | |
| USMT (this work) w/ forward translationUSMT Tuning=Supervised2018.10 | 22.1 | — | — | — | — | |
| Lample et al. (2018b) UNMTUSMT Tuning=No2018.10 | 21 | — | — | — | — | |
| NMT2019.01 | 21 | — | — | — | — | |
| USMT (this work) w/ back-translationUSMT Tuning=Supervised2018.10 | 20.5 | — | — | — | — | |
| USMT (this work) w/ forward translationUSMT Tuning=No2018.10 | 20.2 | — | — | — | — | |
| USMT (this work) w/ back-translationUSMT Tuning=No2018.10 | 19.5 | — | — | — | — | |
| XLMEncoder Pre-training=-, Decoder Pre-training=-2019.01 | 15.3 | — | — | — | — | |
| Unsupervised Neural Machine Translation with Weight SharingWeight-sharing layers=12018.04 | 14.62 | — | — | — | — | |
| Lample et al.2018.04 | 13.33 | — | — | — | — | |
| Word-by-wordMethod type=Baseline2018.04 | 9.34 | — | — | — | — | |
| GPT-3.5-turborole=Victim Model2024.09 | — | 66.1 | 37.7 | 65.2 | 96.5 | |
| Llama3-8Brole=Local Model2024.09 | — | 27.6 | 13 | 35.9 | 87.7 | |
| LoRDMEA method=Locality Reinforced Distillation, query samples=16, initial local model=Llama3-8B2024.09 | — | 58.7 | 30.8 | 58.9 | 91.7 | |
| MLEMEA method=Maximum Likelihood Estimation, query samples=16, initial local model=Llama3-8B2024.09 | — | 57.8 | 30.2 | 57.3 | 90.4 |