Visual Speech Recognition on CMU-MOSEAS-Portuguese (CMpt) (test)
51.6Mean WERVSR model with prediction-based auxiliary tasks
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3+AVSpeech, Training Set=CMpt+MTpt, Training Sets Total Size (hours)=15582022.02 | 51.6 | 0.2 | 51.4 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3, Training Set=CMpt+MTpt, Training Sets Total Size (hours)=9172022.02 | 53.1 | 0.2 | 52.8 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW, Training Set=CMpt+MTpt, Training Sets Total Size (hours)=2562022.02 | 57.2 | 0.7 | 56.6 | |
| CM-seq2seqPre-training Set=LRW, Training Set=CMpt+MTpt, Training Sets Total Size (hours)=2562022.02 | 65.7 | 0.5 | 65.4 |