Speech Translation on MuST-C EN-DE (test-COMMON)
31.7BLEUCascaded
Evaluation Results
| Method | Links | |
|---|---|---|
| Cascaded2022.03 | 31.7 | |
| W2V2-Transformer2022.03 | 31.7 | |
| STEMM2022.03 | 31.7 | |
| STEMM2022.03 | 28.7 | |
| Cascaded2022.03 | 27.5 | |
| AIPABackbone=Enc-Dec, Augmentation=SpecAugment, Objective=All COS2024.06 | 27.5 | |
| AIPABackbone=Enc-Dec, Augmentation=SpecAugment, Auxiliary=PAE2024.06 | 27.39 | |
| W2V2-Transformer2022.03 | 26.9 | |
| AIPABackbone=Enc-Dec, Augmentation=SpecAugment, Objective=Both COS2024.06 | 26.88 | |
| JT Proposed#pars(m)=76, Cross-Attentive Regularization=true, Online Knowledge Distillation=true2021.07 | 26.8 | |
| AIPABackbone=Enc-Dec, Augmentation=SpecAugment, Objective=CTC COS2024.06 | 26.75 | |
| AIPABackbone=Enc-Dec, Augmentation=SpecAugment, Auxiliary=InterCTC2024.06 | 26.68 | |
| AIPABackbone=Enc-Dec, Augmentation=SpecAugment, Objective=CE COS2024.06 | 26.64 | |
| BaselineBackbone=Enc-Dec, Augmentation=SpecAugment, Auxiliary=PAE2024.06 | 26.62 | |
| BaselineBackbone=Enc-Dec, Augmentation=SpecAugment, Auxiliary=InterCTC2024.06 | 26.56 | |
| AIPABackbone=Enc-Dec, Augmentation=SpecAugment2024.06 | 26.38 | |
| BaselineBackbone=Enc-Dec, Augmentation=SpecAugment2024.06 | 26.31 | |
| Pino et al.#pars(m)=4352021.07 | 25.2 | |
| JT-S-MT + CARCross-Attentive Regularization=true2021.07 | 25 | |
| MoSSTMode=Non-streaming2021.09 | 24.9 | |
| JT-S-MT2021.07 | 24.7 | |
| JT-S-ASR2021.07 | 24.4 | |
| JT#pars(m)=762021.07 | 24.1 | |
| Dual-Decoder Transformer (BL)Mode=Non-streaming2021.09 | 23.6 | |
| W-TransfMode=Non-streaming2021.09 | 23.6 | |
| SpeechformerInference Time=1.3x2021.09 | 23.6 | |
| Plain ConvAttentionInference Time=1.8x2021.09 | 23.2 | |
| STASTMode=Non-streaming2021.09 | 23.1 | |
| RealTranSMode=Non-streaming2021.09 | 22.99 | |
| Inaguma et al.2021.07 | 22.9 | |
| Transformer ST ESPnetMode=Non-streaming2021.09 | 22.9 | |
| Inaguma et al.2021.09 | 22.9 | |
| Transformer ST NeurSTMode=Non-streaming2021.09 | 22.8 | |
| Our baselineInference Time=1.0x2021.09 | 22.8 | |
| + compressionInference Time=0.9x2021.09 | 22.8 | |
| Transformer ST FairseqMode=Non-streaming2021.09 | 22.7 | |
| Wang et al.2021.09 | 22.7 | |
| AFS STMode=Non-streaming2021.09 | 22.4 | |
| Wav2Vec2 + TransformerMode=Non-streaming2021.09 | 22.3 | |
| ST#pars(m)=762021.07 | 21.5 | |
| Gangi et al.#pars(m)=302021.07 | 17.7 |