Speech-to-Speech Translation on MuAViC X-to-English
28.7ASR-BLEU (Es->En)AVSR + NMT + TTS + TFG
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| AVSR + NMT + TTS + TFGSystem Type=4-Stage Cascaded, Input-Output Modality=AV-AV, Translation System=NMT2023.12 | 28.7 | 29.21 | 24.54 | 26.3 | |
| ASR + NMT + TTS + TFGSystem Type=4-Stage Cascaded, Input-Output Modality=A-AV, Translation System=NMT2023.12 | 28.66 | 30.55 | 23.54 | 26.14 | |
| Proposed Method (AV2AV)System Type=Direct (Textless), Input-Output Modality=AV-AV, Translation System=AV2AV2023.12 | 26.57 | 31.27 | 23.24 | 24.51 | |
| A2A + TFGSystem Type=2-Stage Cascaded (Textless), Input-Output Modality=A-AV, Translation System=A2A2023.12 | 26.15 | 30.14 | 22.41 | 23.37 | |
| Proposed Method (A2AV)System Type=Direct (Textless), Input-Output Modality=A-AV, Translation System=A2AV2023.12 | 26.04 | 31 | 22.56 | 24.38 | |
| AV2T + TTS + TFGSystem Type=3-Stage Cascaded, Input-Output Modality=AV-AV, Translation System=AV2T2023.12 | 24.61 | 26.9 | 22.33 | 24.83 | |
| A2T + TTS + TFGSystem Type=3-Stage Cascaded, Input-Output Modality=A-AV, Translation System=A2T2023.12 | 24.06 | 27.01 | 21.92 | 24.11 | |
| Proposed Method (V2AV)System Type=Direct (Textless), Input-Output Modality=V-AV, Translation System=V2AV2023.12 | 10.36 | 6.29 | 8.02 | 5.71 |