Multimodal Machine Translation (English-German) on Multi30K 2016 (test)
42.7BLEUVALHALLA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VALHALLATranslation Setting=Text-Only, Ground-truth Visual Tokens usage=false, Model Averaging=true2022.05 | 42.7 | 69.3 | |
| VALHALLA (M)Translation Setting=Multimodal, Ground-truth Visual Tokens usage=true, Model Averaging=true2022.05 | 42.6 | 69.3 | |
| Selective Attn + CATRFeature=CATR2022.03 | 42.5 | 68.81 | |
| Selective Attn + DETRFeature=DETR2022.03 | 42.23 | 68.94 | |
| Gated FusionTranslation Setting=Multimodal2022.05 | 42 | 67.8 | |
| Gated FusionFeature=ResNet2022.03 | 41.96 | 67.84 | |
| Gated FusionSystem Category=Image-must2022.10 | 41.96 | — | |
| Selective Attn + ViT-BaseFeature=ViT-Base2022.03 | 41.93 | 68.55 | |
| Selective Attn + QueryInstFeature=QueryInst2022.03 | 41.9 | 68.64 | |
| VALHALLA (M)Translation Setting=Multimodal, Ground-truth Visual Tokens usage=true, Model Averaging=false2022.05 | 41.9 | 68.7 | |
| VALHALLATranslation Setting=Text-Only, Ground-truth Visual Tokens usage=false, Model Averaging=false2022.05 | 41.9 | 68.8 | |
| Selective AttnFeature=ViT-Large2022.03 | 41.84 | 68.64 | |
| Gated FusionFeature=ViT-Large2022.03 | 41.55 | 68.34 | |
| Doubly-ATTFeature=ResNet2022.03 | 41.45 | 68.04 | |
| RMMTSystem Category=Image-must2022.10 | 41.45 | — | |
| RMMTTranslation Setting=Text-Only2022.05 | 41.4 | 68 | |
| ImaginationFeature=ResNet2022.03 | 41.31 | 68.06 | |
| IKD-MMTSystem Category=Image-free2022.10 | 41.28 | 58.93 | |
| Transformer TinyFeature=Text-only2022.03 | 41.02 | 68.22 | |
| Selective Attn + ViT-SmallFeature=ViT-Small2022.03 | 40.86 | 67.64 | |
| UVR-NMTFeature=ResNet2022.03 | 40.79 | — | |
| Selective Attn + ViT-TinyFeature=ViT-Tiny2022.03 | 40.74 | 67.2 | |
| GMNMTTranslation Setting=Multimodal2022.05 | 39.8 | 57.6 | |
| GMNMTSystem Category=Image-must2022.10 | 39.8 | 57.6 | |
| DCCNTranslation Setting=Multimodal2022.05 | 39.7 | 56.8 | |
| DCCNSystem Category=Image-must2022.10 | 39.7 | 56.8 | |
| CAP-ALLTranslation Setting=Multimodal2022.05 | 39.6 | 57.5 | |
| DS-SUM-L2System Category=Image-must2022.10 | 39.4 | 58.7 | |
| Gumbel-AttentionTranslation Setting=Multimodal2022.05 | 39.2 | 57.8 | |
| Gumbel-Attention MMT2021.03 | 39.2 | 57.8 | |
| Gumbel-attSystem Category=Image-must2022.10 | 39.2 | 57.8 | |
| Multi-modal Transformer2021.03 | 38.7 | 55.7 | |
| MultimodalSystem Category=Image-must2022.10 | 38.7 | 55.7 | |
| ImagiTTranslation Setting=Text-Only2022.05 | 38.5 | 55.7 | |
| ImagiTSystem Category=Image-free2022.10 | 38.5 | 55.7 | |
| Deliberation networks2021.03 | 38 | 55.6 | |
| Del+objSystem Category=Image-must2022.10 | 38 | 55.6 | |
| Text-only Transformer2021.03 | 37.8 | 55.3 | |
| Trg-mul2021.03 | 37.8 | 55.7 | |
| Trg-mulSystem Category=Image-must2022.10 | 37.8 | 57.7 | |
| VMMTFTranslation Setting=Text-Only2022.05 | 37.7 | 56 | |
| Latent Variable MMT2021.03 | 37.7 | 56 | |
| VMMTFSystem Category=Image-free2022.10 | 37.7 | 56 | |
| TransformerSystem Category=Image-free2022.10 | 37.6 | 55.3 | |
| IMGDSystem Category=Image-must2022.10 | 37.3 | 55.1 | |
| Fusion-conv2021.03 | 37 | 57 | |
| Fusion-convSystem Category=Image-must2022.10 | 37 | 57 | |
| UVR-NMTSystem Category=Image-free2022.10 | 36.94 | — | |
| UVR-NMTTranslation Setting=Text-Only2022.05 | 36.9 | — | |
| MultitaskSystem Category=Image-free2022.10 | 36.8 | 55.8 | |
| Doubly-attention2021.03 | 36.5 | 55 | |
| NMTSRC+IMGSystem Category=Image-must2022.10 | 36.5 | 55 |