Machine Translation on Multi30K eng → fra 2017 (test)
62.1BLEUSMT-9B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SMT-9BInput Modality=Text, Speech, Speech Source=Synthetic2026.02 | 62.1 | 89.6 | |
| IMAGEInput Modality=Text, Image, Image Source=Synthetic2026.02 | 61.5 | 86.6 | |
| ConsQA-MMTInput Modality=Text, Image, Image Source=Authentic2026.02 | 58.3 | — | |
| Soul-MixInput Modality=Text, Image, Image Source=Authentic2026.02 | 57.4 | — | |
| D2P-MMT (R)Method Category=Image-dependent Methods, Inference Image Source=Reconstructed, Ensemble=false2025.07 | 56.62 | — | |
| BridgeInput Modality=Text, Image, Image Source=Synthetic2026.02 | 56.2 | — | |
| VALHALLAInput Modality=Text, Image, Image Source=Synthetic2026.02 | 56 | — | |
| VALHALLA*Method Category=Image-free Methods, Ensemble=true2025.07 | 56 | — | |
| VALHALLAMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 56 | — | |
| RG-MMT-EDCInput Modality=Text, Image, Image Source=Authentic2026.02 | 55.8 | — | |
| Noise-robustMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 55.48 | — | |
| D2P-MMT (A)Method Category=Image-dependent Methods, Inference Image Source=Authentic, Ensemble=false2025.07 | 55.13 | — | |
| MMT-VQAMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 54.89 | — | |
| Gated Fusion*Method Category=Image-dependent Methods, Ensemble=true2025.07 | 54.85 | — | |
| IKD-MMTMethod Category=Image-free Methods, Ensemble=false2025.07 | 54.84 | — | |
| NLLB-moe-54BInput Modality=Text2026.02 | 54.8 | 87.7 | |
| Selective AttentionMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 54.52 | — | |
| RMMT*Method Category=Image-free Methods, Ensemble=true2025.07 | 54.4 | — | |
| Transformer-TinyMethod Category=Text-Only Transformer, Ensemble=false2025.07 | 54.35 | — | |
| Gemma3-27B-itInput Modality=Text2026.02 | 54.3 | 87.9 | |
| DCCNMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 54.3 | — | |
| WRA-guidedInput Modality=Text, Image, Image Source=Authentic2026.02 | 54.1 | — | |
| DeepSeek-V3.1Input Modality=Text2026.02 | 54 | 87.7 | |
| Baseline + Lora (Text only)Input Modality=Text, Adaptation=LoRA2026.02 | 54 | 88.2 | |
| Doubly-ATTMethod Category=Image-dependent Methods, Ensemble=false2025.07 | 53.72 | — | |
| ImagiTMethod Category=Image-free Methods, Ensemble=false2025.07 | 52.4 | — | |
| Baseline (Text only)Input Modality=Text2026.02 | 52 | 87.9 | |
| Qwen3-Next-80B-A3BInput Modality=Text2026.02 | 51.9 | 87.6 | |
| UVR-MMTMethod Category=Image-free Methods, Ensemble=false2025.07 | 48.7 | — | |
| DreamLLMInput Modality=Text, Image, Image Source=Synthetic2026.02 | 34.7 | 80.6 |