Speech-to-speech translation on Multi-domain Es->En (CoVOST-2, Europarl-ST, mTEDx)
37.2CoVOST-2 ScoreS2TT-TTS (C2')
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| S2TT-TTS (C2')System category=Cascaded systems, Note=Improved by R-Drop and hyperparameter tuning2022.12 | 37.2 | 34 | 32.5 | 34.6 | |
| S2SpecT2 + t-mBARTSystem category=Direct speech-to-spectrogram systems, Decoder layer configuration=12L→6L, Decoder pre-training=t-mBART2022.12 | 37.2 | 23.7 | 31.7 | 30.9 | |
| S2SpecT2System category=Direct speech-to-spectrogram systems, Decoder layer configuration=6L→6L2022.12 | 37 | 23.4 | 31.3 | 30.6 | |
| UnitY + t-mBARTSystem category=Direct speech-to-unit systems, Decoder layer configuration=12L→2L, Decoder pre-training=t-mBART2022.12 | 36.4 | 33.1 | 32.2 | 33.9 | |
| UnitYSystem category=Direct speech-to-unit systems, Decoder layer configuration=6L→6L2022.12 | 35.4 | 30.8 | 31.3 | 32.5 | |
| S2UT + u-mBART (C5')System category=Direct speech-to-unit systems, Decoder pre-training=u-mBART, Note=Improved by hyperparameter tuning and checkpoint averaging2022.12 | 34.5 | 29.9 | 29.9 | 31.4 | |
| ASR→MT→TTSSystem category=Cascaded systems2022.12 | 33.8 | 29.1 | 32.4 | 31.5 | |
| S2UT + u-mBARTSystem category=Direct speech-to-unit systems, Decoder pre-training=u-mBART2022.12 | 33.5 | 28.6 | 29.1 | 30.4 | |
| ASR→MT→TTS (C1')System category=Cascaded systems, Note=Improved by R-Drop and hyperparameter tuning2022.12 | 32.9 | 34.2 | 30.3 | 32.5 | |
| S2TT-TTSSystem category=Cascaded systems2022.12 | 28.4 | 23.6 | 21.5 | 24.5 |