Visual Speech Recognition on CMU-MOSEAS French
59.1Mean WERVSR model with prediction-based auxiliary tasks
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3+AVSpeech, Training Set=CMfr+MTfr, Training Sets Total Size (hours)=15592022.02 | 59.1 | 0.5 | 58.3 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3, Training Set=CMfr+MTfr, Training Sets Total Size (hours)=9182022.02 | 60.1 | 0.3 | 59.5 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW, Training Set=CMfr+MTfr, Training Sets Total Size (hours)=2572022.02 | 68.4 | 0.5 | 67.5 | |
| CM-seq2seqPre-training Set=LRW, Training Set=CMfr+MTfr, Training Sets Total Size (hours)=2572022.02 | 79.9 | 0.4 | 79.6 |