Multimodal Machine Translation on TopicVD (test)
30.47BLEUProposed Method
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Proposed MethodRetrieval size P=10, Context window w=2, Scaling factor gamma=0.1, Selection count K=5, Fusion weight lambda=0.12026.04 | 30.47 | 38.96 | |
| Segment-Level Video-MMTInput=temporally aligned video segment, Context=local2026.04 | 29.23 | 38.65 | |
| Image-MMTVisual features=extracted from sampled frames, Retrieval=FAISS2026.04 | 28.93 | 38.59 | |
| BigVideoLearning=cross-modal contrastive learning2026.04 | 25.92 | 36.54 | |
| Text-only NMTArchitecture=Transformer-based2026.04 | 19.14 | 33.95 |