Machine Translation on WMT Ro-En 2016 (test)
37.8BLEUmBART
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| mBARTTraining Size=597K, Pre-trained=true2022.03 | 37.8 | — | — | |
| CeMATTraining Size=597K, Pre-trained=true2022.03 | 37.1 | — | — | |
| mRASPTraining Size=597K, Pre-trained=true2022.03 | 36.9 | — | — | |
| XLMTraining Size=597K, Pre-trained=true2022.03 | 35.6 | — | — | |
| Topdown2026.03 | 34.17 | — | — | |
| GLAT+CTCBidirectional=false, Parameters=62M x 2, Speedup=16.8x, Decoding=greedy2021.05 | 34.16 | — | — | |
| Transformerimplementation=ours2020.08 | 34.05 | — | 1 | |
| duplex REDER (final model)Bidirectional=true, Parameters=58M, Speedup=5.5x, Decoding=beam search (size 20)2021.05 | 34.03 | — | — | |
| DirectTraining Size=597K, Pre-trained=false2022.03 | 34 | — | — | |
| ATIterations=N, latency (ms)=4862020.11 | 33.98 | — | — | |
| Transformer2026.03 | 33.93 | — | — | |
| MGNMTBidirectional=true, Parameters=195M2021.05 | 33.9 | — | — | |
| GLAT + CTCIdec=12020.08 | 33.84 | — | 14.6 | |
| Transformer-base (KD teacher)Bidirectional=false, Parameters=62M x 2, Speedup=1.0x2021.05 | 33.7 | — | — | |
| GLAT + NPDIdec=1, m=72020.08 | 33.51 | — | 7.9 | |
| GLATBidirectional=false, Parameters=62M x 2, Speedup=15.3x, Decoding=greedy2021.05 | 33.51 | — | — | |
| CTC*Idec=12020.08 | 33.46 | — | 14.6 | |
| CMLMIterations=10, latency (ms)=1662020.11 | 33.31 | — | — | |
| Mask-PredictIdec=10, type=Iterative NAT2020.08 | 33.31 | — | 1.7 | |
| LATIterations=4, latency (ms)=732020.11 | 33.26 | — | — | |
| LevTIdec=6+2020.08 | 33.26 | — | 4 | |
| CMLMIterations=4, latency (ms)=722020.11 | 33.23 | — | — | |
| 1.5-entmaxactivation=1.5-entmax2019.08 | 33.1 | — | — | |
| simplex REDERBidirectional=false, Parameters=58M x 2, Speedup=5.5x, Decoding=beam search (size 20)2021.05 | 32.98 | — | — | |
| alpha-entmaxactivation=alpha-entmax2019.08 | 32.89 | — | — | |
| Flowseq + NPDIdec=1, m=302020.08 | 32.84 | — | — | |
| softmaxactivation=softmax2019.08 | 32.7 | — | — | |
| GLATIdec=12020.08 | 32.04 | — | 15.3 | |
| imit-NAT + NPDIdec=1, m=72020.08 | 31.81 | — | 9.7 | |
| XLMEncoder Pre-training=MLM, Decoder Pre-training=MLM2019.01 | 31.8 | — | — | |
| Autoregressivebeam size=42017.11 | 31.76 | 607 | 1 | |
| ImputerIdec=12020.08 | 31.7 | — | 18.6 | |
| ImputerBidirectional=false, Parameters=58M x 2, Speedup=18.6x, Decoding=greedy2021.05 | 31.7 | — | — | |
| CTCBidirectional=false, Parameters=58M x 2, Speedup=18.6x, Decoding=greedy2021.05 | 31.6 | — | — | |
| NATdecoding=Noisy Parallel Decoding, fine-tuning=true, sample size=1002017.11 | 31.44 | 257 | 2.36 | |
| NAT-FT + NPDIdec=1, m=1002020.08 | 31.44 | — | 2.4 | |
| LATIterations=1, latency (ms)=312020.11 | 31.24 | — | — | |
| Autoregressivebeam size=12017.11 | 31.03 | 408 | 1.49 | |
| CONTBackbone=T5-small, Loss Function=N-Pairs loss, target-source representation similarity=true2022.05 | 30.91 | — | — | |
| NATdecoding=Noisy Parallel Decoding, fine-tuning=true, sample size=102017.11 | 30.76 | 79 | 7.68 | |
| FlowseqIdec=12020.08 | 30.72 | — | 1.1 | |
| CONTBackbone=T5-small, Loss Function=N-Pairs loss, target-source representation similarity=false2022.05 | 30.54 | — | — | |
| XLMEncoder Pre-training=MLM, Decoder Pre-training=-2019.01 | 30.5 | — | — | |
| XLMEncoder Pre-training=MLM, Decoder Pre-training=CLM2019.01 | 30.4 | — | — | |
| NAT-IRIdec=102020.08 | 30.19 | — | 1.5 | |
| Naive CLBackbone=T5-small, Loss Function=N-Pairs loss, target-source representation similarity=true2022.05 | 29.86 | — | — | |
| XLMEncoder Pre-training=CLM, Decoder Pre-training=MLM2019.01 | 29.8 | — | — | |
| Naive CLBackbone=T5-small, Loss Function=N-Pairs loss, target-source representation similarity=false2022.05 | 29.74 | — | — | |
| CONTBackbone=T5-small, Loss Function=InfoNCE loss2022.05 | 29.64 | — | — | |
| NAT-base*Idec=12020.08 | 29.43 | — | 15.3 | |
| CLAPSBackbone=T5-small, Loss Function=InfoNCE loss2022.05 | 29.41 | — | — | |
| LaNMTIdec=42020.08 | 29.1 | — | 5.7 | |
| Dropout CLBackbone=T5-small, Loss Function=InfoNCE loss2022.05 | 29.1 | — | — | |
| NATdecoding=argmax, fine-tuning=true2017.11 | 29.06 | 39 | 15.6 | |
| NAT-FTIdec=12020.08 | 29.06 | — | 15.6 | |
| vanilla NATBidirectional=false, Parameters=62M x 2, Speedup=15.6x, Decoding=greedy2021.05 | 29.06 | — | — | |
| imit-NATIdec=12020.08 | 28.9 | — | 18.6 | |
| SSMBA CLBackbone=T5-small, Loss Function=InfoNCE loss2022.05 | 28.48 | — | — | |
| MLEBackbone=T5-small, Loss Function=MLE2022.05 | 28.21 | — | — | |
| CMLMIterations=1, latency (ms)=272020.11 | 28.2 | — | — | |
| Mask-PredictIdec=1, type=Fully NAT2020.08 | 28.2 | — | — | |
| XLMEncoder Pre-training=-, Decoder Pre-training=CLM2019.01 | 28 | — | — | |
| NATdecoding=argmax2017.11 | 27.83 | 39 | 15.6 | |
| XLMEncoder Pre-training=CLM, Decoder Pre-training=CLM2019.01 | 27.8 | — | — | |
| Naive CLBackbone=T5-small, Loss Function=InfoNCE loss2022.05 | 27.79 | — | — | |
| CONTBackbone=Transformer-small, Loss Function=N-Pairs loss, target-source representation similarity=true2022.05 | 27.7 | — | — | |
| CONTBackbone=Transformer-small, Loss Function=N-Pairs loss, target-source representation similarity=false2022.05 | 27.42 | — | — | |
| XLMEncoder Pre-training=MLM, Decoder Pre-training=-2019.01 | 27.3 | — | — | |
| XLMEncoder Pre-training=EMB, Decoder Pre-training=EMB2019.01 | 26.6 | — | — | |
| Naive CLBackbone=Transformer-small, Loss Function=N-Pairs loss, target-source representation similarity=false2022.05 | 26.27 | — | — | |
| Naive CLBackbone=Transformer-small, Loss Function=N-Pairs loss, target-source representation similarity=true2022.05 | 26.15 | — | — | |
| Dropout CLBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 26.01 | — | — | |
| SSMBA CLBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 25.98 | — | — | |
| MLEBackbone=Transformer-small, Loss Function=MLE2022.05 | 25.78 | — | — | |
| CONTBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 25.74 | — | — | |
| Naive CLBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 25.49 | — | — | |
| NAT-CTCIdec=12020.08 | 24.67 | — | — | |
| CTC w/o KDBidirectional=false, Parameters=58M x 2, Decoding=greedy2021.05 | 24.67 | — | — | |
| XLMEncoder Pre-training=CLM, Decoder Pre-training=-2019.01 | 24.6 | — | — | |
| PBSMT + NMT2019.01 | 23.9 | — | — | |
| CLAPSBackbone=Transformer-small, Loss Function=InfoNCE loss2022.05 | 23.59 | — | — | |
| PBSMT2019.01 | 23 | — | — | |
| NMT2019.01 | 19.4 | — | — | |
| XLMEncoder Pre-training=-, Decoder Pre-training=-2019.01 | 18.3 | — | — |