Visual Speech Recognition on Multilingual TEDx Italian (MTit) (test)
57.9Mean WERVSR model with prediction-based auxiliary tasks
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3+AVSpeech, Training Set=MTit, Training Sets Total Size (hours)=15052022.02 | 57.9 | 0.7 | 57.4 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3, Training Set=MTit, Training Sets Total Size (hours)=8642022.02 | 58.7 | 0.3 | 58.2 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW, Training Set=MTit, Training Sets Total Size (hours)=2032022.02 | 65.9 | 0.5 | 65.2 | |
| CM-seq2seqPre-training Set=LRW, Training Set=MTit, Training Sets Total Size (hours)=2032022.02 | 71.5 | 0.4 | 70.9 |