Multimodal Machine Translation on Multi30K (test)
60.6BLEU-4MMT
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| MMTModel Type=Multimodal Machine Translation2019.06 | 60.6 | 75 | — | — | |
| delModel Type=Deliberation Network, Image Information=None2019.06 | 60.1 | 74.6 | — | — | |
| ImagiTInput Modality=Multimodal, ground truth=true2020.09 | 59.9 | 74.3 | — | — | |
| del+objModel Type=Deliberation Network, Image Information=AIF object (del+obj)2019.06 | 59.8 | 74.4 | — | — | |
| Transformer+AttInput Modality=Multimodal2020.09 | 59.8 | 74.4 | — | — | |
| ImagiTInput Modality=Text-only2020.09 | 59.7 | 74 | — | — | |
| del+sumModel Type=Deliberation Network, Image Information=AIC (del+sum)2019.06 | 59.6 | 74.3 | — | — | |
| base+sumModel Type=Transformer (base), Image Information=AIC (base+sum)2019.06 | 59.2 | 73.9 | — | — | |
| del+attModel Type=Deliberation Network, Image Information=AIF spatial (del+att)2019.06 | 59.2 | 73.7 | — | — | |
| baseModel Type=Transformer (base), Image Information=None2019.06 | 59 | 73.7 | — | — | |
| TransformerInput Modality=Text-only2020.09 | 59 | 73.6 | — | — | |
| base+attModel Type=Transformer (base), Image Information=AIF spatial (base+att)2019.06 | 58.7 | 73.5 | — | — | |
| Lookup tableInput Modality=Text-only2020.09 | 57.5 | — | — | — | |
| base+objModel Type=Transformer (base), Image Information=AIF object (base+obj)2019.06 | 57.3 | 72.9 | — | — | |
| trg-mulInput Modality=Multimodal2020.09 | 54.7 | 71.3 | — | — | |
| DCCN#Params=+0.9M2020.09 | 54.3 | 70.3 | — | — | |
| VAG-NMTModality=Multimodal2018.08 | 53.8 | 70.3 | — | — | |
| Encoder-attention#Params=+1.0M2020.09 | 53.8 | 69.9 | — | — | |
| VAG-NMTInput Modality=Multimodal2020.09 | 53.8 | 70.3 | — | — | |
| Doubly-attention#Params=+3.9M2020.09 | 53.6 | 70 | — | — | |
| Text-Only NMTModality=Text-only2018.08 | 53.5 | 70 | — | — | |
| fusion-convInput Modality=Multimodal2020.09 | 53.5 | 70.4 | — | — | |
| Transformer#Params=16.0M2020.09 | 53.3 | 69.7 | — | — | |
| ImagiTInput Modality=Multimodal, ground truth=true2020.09 | 52.8 | 68.6 | — | — | |
| LIUMCVCModality=Multimodal2018.08 | 52.7 | 69.5 | — | — | |
| Trg-mul#Params=-2020.09 | 52.7 | 69.5 | — | — | |
| trg-mulInput Modality=Multimodal2020.09 | 52.7 | 69.5 | — | — | |
| ImagiTInput Modality=Text-only2020.09 | 52.4 | 68.3 | — | — | |
| TransformerInput Modality=Text-only2020.09 | 51.9 | 68.3 | — | — | |
| Fusion-conv#Params=-2020.09 | 51.6 | 68.6 | — | — | |
| fusion-convInput Modality=Multimodal2020.09 | 51.6 | 68.6 | — | — | |
| Lookup tableInput Modality=Text-only2020.09 | 48.5 | — | — | — | |
| MeMADData Augmentation=true, Ensemble=true2021.02 | 45.5 | — | — | — | |
| VL-T5V&L PT=true2021.02 | 45.5 | — | — | — | |
| VL-T5V&L PT=false2021.02 | 45.3 | — | — | — | |
| MeMADData Augmentation=true2021.02 | 45.1 | — | — | — | |
| T5Text-only=true2021.02 | 44.6 | — | — | — | |
| VL-T5V&L PT=false2021.02 | 42.4 | — | — | — | |
| MeMADData Augmentation=true, Ensemble=true2021.02 | 41.8 | — | — | — | |
| T5Text-only=true2021.02 | 41.6 | — | — | — | |
| VL-BARTV&L PT=false2021.02 | 41.3 | — | — | — | |
| BARTText-only=true2021.02 | 41.2 | — | — | — | |
| VL-T5V&L PT=true2021.02 | 40.9 | — | — | — | |
| MeMADData Augmentation=true2021.02 | 40.8 | — | — | — | |
| DCCN#Params=+1.0M2020.09 | 39.7 | 56.8 | — | — | |
| MSAData Augmentation=true2021.02 | 39.5 | — | — | — | |
| Encoder-attention#Params=+1.1M2020.09 | 39 | 56.6 | — | — | |
| MeMAD2021.02 | 38.9 | — | — | — | |
| Doubly-attention#Params=+4.0M2020.09 | 38.7 | 56.4 | — | — | |
| MultimodalInput Modality=Multimodal2020.09 | 38.7 | 55.7 | — | — | |
| MSA2021.02 | 38.7 | — | — | — | |
| ImagiTInput Modality=Multimodal, ground truth=true2020.09 | 38.6 | 55.7 | — | — | |
| ImagiTInput Modality=Text-only2020.09 | 38.5 | 55.7 | — | — | |
| VMMTFTraining Sentences=145K, Prior Type=Fixed2018.11 | 38.4 | 56 | — | — | |
| VMMTCTraining Sentences=145K, Prior Type=Conditional2018.11 | 38.4 | 56.3 | — | — | |
| MMTModel Type=Multimodal Machine Translation2019.06 | 38.4 | 53.1 | — | — | |
| Transformer#Params=16.1M2020.09 | 38.4 | 56 | — | — | |
| Stochastic attention2020.09 | 38.2 | 55.4 | — | — | |
| del+objModel Type=Deliberation Network, Image Information=AIF object (del+obj)2019.06 | 38 | 55.6 | — | — | |
| Deliberation Network2020.09 | 38 | 55.6 | — | — | |
| Transformer+AttInput Modality=Multimodal2020.09 | 38 | 55.6 | — | — | |
| ImaginationTraining Sentences=654K2018.11 | 37.8 | 57.1 | — | — | |
| Trg-mul2020.09 | 37.8 | 57.7 | — | — | |
| trg-mulInput Modality=Multimodal2020.09 | 37.8 | 57.7 | — | — | |
| NMT2018.11 | 37.7 | 56 | — | — | |
| delModel Type=Deliberation Network, Image Information=None2019.06 | 37.7 | 55.5 | — | — | |
| Latent Variable MMT2020.09 | 37.7 | 56 | — | — | |
| VL-BARTV&L PT=true2021.02 | 37.7 | — | — | — | |
| TransformerInput Modality=Text-only2020.09 | 37.6 | 55.3 | — | — | |
| VMMTFInput Modality=Text-only2020.09 | 37.6 | 56 | — | — | |
| IMGDvisual_feature_integration=image-initialised decoder2017.01 | 37.3 | 55.1 | 42.8 | 67.7 | |
| del+sumModel Type=Deliberation Network, Image Information=AIC (del+sum)2019.06 | 37.3 | 55.2 | — | — | |
| IMG_DInput Modality=Multimodal2020.09 | 37.3 | 55.1 | — | — | |
| del+attModel Type=Deliberation Network, Image Information=AIF spatial (del+att)2019.06 | 37.2 | 55.1 | — | — | |
| IMG1wvisual_feature_integration=images as words in the source sentence2017.01 | 37.1 | 54.5 | 42.7 | 66.9 | |
| IMGEvisual_feature_integration=image-initialised encoder2017.01 | 37.1 | 55 | 43.1 | 67.6 | |
| IMGE+Dvisual_feature_integration=image-initialised encoder + image-initialised decoder2017.01 | 37 | 54.7 | 42.6 | 67.2 | |
| Fusion-conv2020.09 | 37 | 57 | — | — | |
| fusion-convInput Modality=Multimodal2020.09 | 37 | 57 | — | — | |
| IMG2wvisual_feature_integration=images as words in the source sentence2017.01 | 36.9 | 54.3 | 41.9 | 66.8 | |
| base+attModel Type=Transformer (base), Image Information=AIF spatial (base+att)2019.06 | 36.9 | 54.5 | — | — | |
| Lookup tableInput Modality=Text-only2020.09 | 36.9 | — | — | — | |
| Imagination2020.09 | 36.8 | 55.8 | — | — | |
| MultitaskInput Modality=Text-only2020.09 | 36.8 | 55.8 | — | — | |
| Huang + RCNNmodel_id=m3, multimodal_strategy=additional object detections2017.01 | 36.5 | 54.1 | — | — | |
| NMT_SRC+IMGInput Modality=Multimodal2020.09 | 36.5 | 55 | — | — | |
| baseModel Type=Transformer (base), Image Information=None2019.06 | 36.4 | 54.5 | — | — | |
| base+objModel Type=Transformer (base), Image Information=AIF object (base+obj)2019.06 | 36.4 | 54.5 | — | — | |
| base+sumModel Type=Transformer (base), Image Information=AIC (base+sum)2019.06 | 35.9 | 54.2 | — | — | |
| VL-BARTV&L PT=false2021.02 | 35.9 | — | — | — | |
| IMG2w+Dvisual_feature_integration=images as words + image-initialised decoder2017.01 | 35.7 | 53.6 | 43.3 | 66.2 | |
| BARTText-only=true2021.02 | 35.4 | — | — | — | |
| Huangmodel_id=m1, multimodal_strategy=image at head2017.01 | 35.1 | 52.2 | — | — | |
| NMTtype=text-only baseline2017.01 | 33.7 | 52.3 | 46.7 | 64.5 | |
| PBSMTtype=text-only baseline2017.01 | 32.9 | 54.1 | 45.1 | 67.4 | |
| ImagiTInput Modality=Multimodal, ground truth=true2020.09 | 32.4 | 52.5 | — | — | |
| ImagiTInput Modality=Text-only2020.09 | 32.1 | 52.4 | — | — | |
| MeMAD2021.02 | 32 | — | — | — | |
| TransformerInput Modality=Text-only2020.09 | 31.7 | 52.1 | — | — | |
| Text-Only NMTModality=Text-only2018.08 | 31.6 | 52.2 | — | — |