Speech Recognition on German video dataset (noisy)
12.3WER+ ce pretrain + ce loss, mid, top
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| + ce pretrain + ce loss, mid, toplambda_ce=0.62020.11 | 12.3 | 4.4 | |
| + crosslingual pretrain + aux + kl losslambda_aux=0.32020.11 | 12.4 | 3.6 | |
| + ce pretrain + ce loss, midlambda_ce=0.62020.11 | 12.5 | 3.2 | |
| + aux losslambda_aux=0.32020.11 | 12.6 | 2.8 | |
| + aux + kl losslambda_aux=0.32020.11 | 12.6 | 2.8 | |
| + kl losslambda_aux=0.32020.11 | 12.8 | 1.2 | |
| + crosslingual pretrain2020.11 | 12.8 | 1.6 | |
| + ce pretrain2020.11 | 12.8 | 1.2 | |
| baseline2020.11 | 13 | — |