Video Speech Recognition on Multilingual TEDx-French (MTfr) (test)
67Mean WERVSR model with prediction-based auxiliary tasks
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3, Training Set=CMfr+MTfr, Training Sets Total Size (hours)=9182022.02 | 67 | 30 | 66.7 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3+AVSpeech, Training Set=CMfr+MTfr, Training Sets Total Size (hours)=15592022.02 | 67 | 60 | 66.2 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW, Training Set=CMfr+MTfr, Training Sets Total Size (hours)=2572022.02 | 74.6 | 60 | 73.4 | |
| CM-seq2seqPre-training Set=LRW, Training Set=CMfr+MTfr, Training Sets Total Size (hours)=2572022.02 | 84 | 70 | 83.2 |