Speech-to-speech translation on CVSS-C
0.911Avg ScoreSynthetic target
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Synthetic targettype=Baseline/Reference2022.12 | 0.911 | 0.884 | 0.895 | 0.93 | |
| UnitY + w2v-BERT + t-mBARTCategory=Direct speech-to-unit systems, Encoder=w2v-BERT, Decoder=t-mBART2022.12 | 0.474 | 0.607 | 0.521 | 0.41 | |
| S2UT + w2v-BERT + u-mBARTCategory=Direct speech-to-unit systems, Encoder=w2v-BERT, Decoder=u-mBART2022.12 | 0.445 | 0.588 | 0.495 | 0.377 | |
| S2TT + TTS + w2v-BERT + t-mBARTCategory=Cascaded systems, Encoder=w2v-BERT, Decoder=t-mBART2022.12 | 0.42 | 0.533 | 0.463 | 0.365 | |
| S2SpecT2 + w2v-BERT + t-mBARTCategory=Direct speech-to-spectrogram systems, Encoder=w2v-BERT, Decoder=t-mBART2022.12 | 0.419 | 0.592 | 0.492 | 0.331 | |
| S2SpecT + w2v-BERTCategory=Direct speech-to-spectrogram systems, Encoder=w2v-BERT2022.12 | 0.395 | 0.582 | 0.461 | 0.306 | |
| S2SpecT2 + S2TT pre-trainingCategory=Direct speech-to-spectrogram systems, Pre-training=S2TT2022.12 | 0.336 | 0.566 | 0.417 | 0.226 | |
| UnitY + S2TT pre-trainingCategory=Direct speech-to-unit systems, Pre-training=S2TT2022.12 | 0.333 | 0.572 | 0.415 | 0.22 | |
| S2UT + S2TT pre-trainingCategory=Direct speech-to-unit systems, Pre-training=S2TT2022.12 | 0.329 | 0.55 | 0.405 | 0.224 | |
| UnitYCategory=Direct speech-to-unit systems2022.12 | 0.312 | 0.564 | 0.396 | 0.192 | |
| S2SpecT + S2TT pre-trainingCategory=Direct speech-to-spectrogram systems, Pre-training=S2TT2022.12 | 0.311 | 0.521 | 0.377 | 0.213 | |
| S2SpecT2Category=Direct speech-to-spectrogram systems2022.12 | 0.306 | 0.56 | 0.389 | 0.187 | |
| S2TT + TTSCategory=Cascaded systems2022.12 | 0.304 | 0.504 | 0.384 | 0.204 | |
| S2UTCategory=Direct speech-to-unit systems2022.12 | 0.294 | 0.536 | 0.356 | 0.188 | |
| S2SpecTCategory=Direct speech-to-spectrogram systems2022.12 | 0.273 | 0.498 | 0.328 | 0.175 | |
| UnitYSystem Type=Direct speech-to-unit, Pre-training=w2v-BERT + t-mBART2022.12 | 0.245 | 0.346 | 0.289 | 0.193 | |
| Translatotron2System Type=Direct speech-to-spectrogram, Decoder=Transformer, Backbone=mSLAM, Augmentation=TTS2022.12 | 0.22 | 0.335 | 0.258 | 0.165 | |
| S2UTSystem Type=Direct speech-to-unit, Pre-training=w2v-BERT + u-mBART2022.12 | 0.208 | 0.316 | 0.254 | 0.154 | |
| Translatotron2System Type=Direct speech-to-spectrogram, Decoder=Transformer, Backbone=mSLAM2022.12 | 0.193 | 0.332 | 0.246 | 0.125 | |
| S2SpecT2System Type=Direct speech-to-spectrogram, Pre-training=w2v-BERT + t-mBART2022.12 | 0.186 | 0.321 | 0.247 | 0.116 | |
| Translatotron2System Type=Direct speech-to-spectrogram, Decoder=Transformer, Backbone=w2v-BERT2022.12 | 0.179 | 0.325 | 0.229 | 0.109 | |
| S2SpecTSystem Type=Direct speech-to-spectrogram, Backbone=w2v-BERT2022.12 | 0.166 | 0.305 | 0.219 | 0.098 | |
| S2TT + TTSSystem Type=Cascaded systems, Pre-training=w2v-BERT + t-mBART2022.12 | 0.149 | 0.211 | 0.182 | 0.115 | |
| S2SpecT2System Type=Direct speech-to-spectrogram, S2TT pre-training=true2022.12 | 0.131 | 0.298 | 0.188 | 0.052 | |
| UnitYSystem Type=Direct speech-to-unit, S2TT pre-training=true2022.12 | 0.13 | 0.304 | 0.187 | 0.048 | |
| S2TT -> TTSSystem Type=Cascaded systems, ASR pre-training=true2022.12 | 0.127 | 0.307 | 0.183 | 0.044 | |
| Translatotron2System Type=Direct speech-to-spectrogram, Decoder=Transformer, S2TT pre-training=true2022.12 | 0.12 | 0.297 | 0.166 | 0.042 | |
| UnitYSystem Type=Direct speech-to-unit2022.12 | 0.12 | 0.29 | 0.178 | 0.04 | |
| S2UTSystem Type=Direct speech-to-unit, S2TT pre-training=true2022.12 | 0.114 | 0.272 | 0.164 | 0.04 | |
| S2SpecT2System Type=Direct speech-to-spectrogram2022.12 | 0.113 | 0.291 | 0.169 | 0.031 | |
| S2TT -> TTSSystem Type=Cascaded systems2022.12 | 0.106 | 0.288 | 0.155 | 0.024 | |
| Translatotron2System Type=Direct speech-to-spectrogram, Decoder=Transformer2022.12 | 0.101 | 0.269 | 0.142 | 0.028 | |
| S2SpecTSystem Type=Direct speech-to-spectrogram, S2TT pre-training=true2022.12 | 0.096 | 0.239 | 0.138 | 0.032 | |
| S2UTSystem Type=Direct speech-to-unit2022.12 | 0.091 | 0.259 | 0.129 | 0.019 | |
| Translatotron2System Type=Direct speech-to-spectrogram2022.12 | 0.087 | 0.254 | 0.126 | 0.015 | |
| S2TT + TTSSystem Type=Cascaded systems2022.12 | 0.078 | 0.182 | 0.119 | 0.026 | |
| S2SpecTSystem Type=Direct speech-to-spectrogram2022.12 | 0.076 | 0.218 | 0.106 | 0.015 | |
| TranslatotronSystem Type=Direct speech-to-spectrogram2022.12 | 0.034 | 0.119 | 0.035 | 0.003 |