Visual Speech Recognition on LSVSR (test)
18.3Word Error RateAudio-Ph
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Audio-PhParams=58M2018.07 | 18.3 | 12.5 | 11.5 | |
| V2PParams=49M, Language Model=5-gram word LM2018.07 | 40.9 | 33.6 | 28.3 | |
| V2P-FullyConvParams=29M, Temporal Aggregation=6 dilated temporal convolution layers2018.07 | 51.6 | 41.3 | 36.7 | |
| V2P-NoLMParams=49M, Language Model=None (uni-gram dictionary)2018.07 | 53.6 | 33.6 | 34.6 | |
| Baseline-LipNet-Large-PhParams=40M, Normalization=None, Dropout=True2018.07 | 72.7 | 53 | 54 | |
| Baseline-Seq2seq-ChParams=15M, Architecture=Sequence-to-sequence (WAS variant)2018.07 | 76.8 | — | 49.9 | |
| Professional w/ context2018.07 | 86.4 | — | — | |
| Baseline-LipNet-PhParams=7M, Prediction Type=Phonemes2018.07 | 89.8 | 65.8 | 72.8 | |
| Professional w/o context2018.07 | 92.9 | — | — | |
| Baseline-LipNet-ChParams=7M, Prediction Type=Characters2018.07 | 93 | — | 64.6 |