Visual Speech Recognition on Multilingual TEDx-Portuguese (MTpt) (test)
70.2Mean AccuracyCM-seq2seq
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CM-seq2seqPre-training Set=LRW, Training Set=CMpt+MTpt, Training Sets Total Size (hours)=2562022.02 | 70.2 | 0.3 | 69.7 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW, Training Set=CMpt+MTpt, Training Sets Total Size (hours)=2562022.02 | 66 | 0.5 | 65.3 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3, Training Set=CMpt+MTpt, Training Sets Total Size (hours)=9172022.02 | 62.4 | 0.4 | 62 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3+AVSpeech, Training Set=CMpt+MTpt, Training Sets Total Size (hours)=15582022.02 | 62.1 | 0.6 | 61.5 |