Machine Translation on MSCOCO En-De
36.2BLEUmBART + MT* w/ adapters
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| mBART + MT* w/ adaptersObjectives=NMT + MLM, trainable_params=12.6M2022.12 | 36.2 | 57.4 | — | |
| VGAMTObjectives=MMT + VMLM, trainable_params=13.2M2022.12 | 35.7 | 54.4 | — | |
| Noise-robustMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 31.09 | — | — | |
| D2P-MMT (R)Method Category=Image-dependent Methods, Inference Image Source=Reconstructed, Ensemble=false2025.07 | 31.01 | — | — | |
| MMT-VQAMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 30.96 | — | — | |
| D2P-MMT (A)Method Category=Image-dependent Methods, Inference Image Source=Authentic, Ensemble=false2025.07 | 30.93 | — | — | |
| VALHALLA*Method Category=Image-free Methods, Ensemble=true2025.07 | 30.7 | — | — | |
| VALHALLAMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 30.7 | — | — | |
| Transformer-TinyMethod Category=Text-Only Transformer, Ensemble=false2025.07 | 30.52 | — | — | |
| Selective AttentionMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 30.22 | — | — | |
| IKD-MMTMethod Category=Image-free Methods, Ensemble=false2025.07 | 30.17 | — | — | |
| RMMT*Method Category=Image-free Methods, Ensemble=true2025.07 | 30 | — | — | |
| Doubly-ATTMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 29.63 | — | — | |
| TLM + MT*Objectives=NMT, trainable_params=42M2022.12 | 29.4 | 15.2 | — | |
| Gated Fusion*Method Category=Image-dependent Methods, Ensemble=true2025.07 | 29.04 | — | — | |
| ImagiTMethod Category=Image-free Methods, Ensemble=false2025.07 | 28.7 | — | — | |
| VTLM + MMT*Objectives=MMT, trainable_params=44M2022.12 | 28.2 | 16.8 | — | |
| Vanilla MT*Objectives=NMT, trainable_params=4.1M2022.12 | 27.8 | 9.2 | — | |
| mBART + MT*Objectives=NMT, trainable_params=-2022.12 | 27.6 | 38.3 | — | |
| Gumbel-AttentionMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 26.9 | — | — | |
| DCCNMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 26.7 | — | — | |
| Gated Fusion*Objectives=MMT, trainable_params=2.8M2022.12 | 26.6 | 5.5 | — | |
| Graph-MMT*Objectives=MMT, trainable_params=4.1M2022.12 | 25.9 | 6 | — | |
| Gated FusionModality=Multimodal2022.12 | — | — | 44.2 | |
| Graph-MMTModality=Multimodal2022.12 | — | — | 43.2 | |
| mBART + MTModality=Text-only, Adapters=false2022.12 | — | — | 44.2 | |
| mBART + MTModality=Text-only, Adapters=true2022.12 | — | — | 51.7 | |
| TLM + MTModality=Text-only2022.12 | — | — | 46.1 | |
| Vanilla MTModality=Text-only2022.12 | — | — | 45.3 | |
| VGAMTModality=Multimodal2022.12 | — | — | 51.7 | |
| VTLM + MMTModality=Multimodal2022.12 | — | — | 45.6 |