Lip Reading on LRS2 (test)
14.6WERAuto-AVSR
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Auto-AVSRPhoneme-based=false, Input=Video, Extra Data=true, Total Hours=34482025.07 | 14.6 | — | — | |
| PV-ASRPhoneme-based=true, Input=Video+117-Points, LLM=NLLB(1.3B), Extra Data=false, Total Hours=2232025.07 | 16 | — | — | |
| V-ASRPhoneme-based=true, Input=Video, LLM=NLLB(1.3B), Extra Data=false, Total Hours=2232025.07 | 17.1 | — | — | |
| VALLRPhoneme-based=true, Input=Video, LLM=LLaMA, Extra Data=false, Total Hours=282025.07 | 20.8 | — | — | |
| Visual Transformer Pooling (VTP)Datasets used (Training)=LRS2, LRS3, MV-LRS†, TEDxext, Total # hours (Training)=2,6762021.10 | 22.6 | — | — | |
| VTPPhoneme-based=false, Input=Video, Extra Data=true, Total Hours=26762025.07 | 22.6 | — | — | |
| OpenSRType=Full-Shot, Labeled Utt(hrs) Video=224, Labeled Utt(hrs) Audio=2242023.06 | 25 | — | — | |
| Auto-AVSRPhoneme-based=false, Input=Video, Extra Data=true, Total Hours=8182025.07 | 27.9 | — | — | |
| Shi et al.Type=Full-Shot, Labeled Utt(hrs) Video=2242023.06 | 28.6 | — | — | |
| Visual Transformer Pooling (VTP)Datasets used (Training)=LRS2, LRS3, Total # hours (Training)=6982021.10 | 28.9 | — | — | |
| Prajwal et al.Type=Full-Shot, Labeled Utt(hrs) Video=6982023.06 | 28.9 | — | — | |
| CM-auxPhoneme-based=false, Input=Video, Extra Data=false, Total Hours=2232025.07 | 32.9 | — | — | |
| OpenSRType=Zero-Shot, Labeled Utt(hrs) Video=X, Labeled Utt(hrs) Audio=2242023.06 | 36 | — | — | |
| Hyb. + ConformerDatasets used (Training)=LRS2, LRW, Total # hours (Training)=3892021.10 | 37.9 | — | — | |
| Hyb.-Conf.Phoneme-based=false, Input=Video, Extra Data=true, Total Hours=3812025.07 | 37.9 | — | — | |
| Ma et al.Type=Few-Shot, Labeled Utt(hrs) Video=2242023.06 | 39.1 | — | — | |
| OpenSRType=Zero-Shot, Labeled Utt(hrs) Video=X, Labeled Utt(hrs) Audio=292023.06 | 39.2 | — | — | |
| Proposed Method (MVM)Backbone=R18 + MS-TCN2022.04 | 44.5 | — | — | |
| TM-seq2seqDatasets used (Training)=LRS2, LRS3, LRW, MV-LRS†, Total # hours (Training)=1,6372021.10 | 48.3 | — | — | |
| TDNNDatasets used (Training)=LRS2, Total # hours (Training)=2242021.10 | 48.9 | — | — | |
| VTPPhoneme-based=false, Input=Video, Extra Data=true, Total Hours=6982025.07 | 48.9 | — | — | |
| Ren et al.Type=Full-Shot, Labeled Utt(hrs) Video=698, Labeled Utt(hrs) Audio=6982023.06 | 49.2 | — | — | |
| TM-seq2seq2019.11 | 49.8 | — | — | |
| Baseline (Afouras et al. 2018)2022.04 | 49.8 | — | — | |
| Afouras et al.Type=Full-Shot, Labeled Utt(hrs) Video=6982023.06 | 49.8 | — | — | |
| CTC + KDDatasets used (Training)=LRS2, LRS3, VoxCeleb2‡, Total # hours (Training)=1,0322021.10 | 51.3 | — | — | |
| Afouras et al.Type=Full-Shot, Labeled Utt(hrs) Video=224, Labeled Utt(hrs) Audio=8082023.06 | 51.3 | — | — | |
| Conv-seq2seqDatasets used (Training)=LRS2, LRS3, Total # hours (Training)=6982021.10 | 51.7 | — | — | |
| Zhang et al.Type=Full-Shot, Labeled Utt(hrs) Video=6982023.06 | 51.7 | — | — | |
| Afouras et al.Type=Few-Shot, Labeled Utt(hrs) Video=224, Labeled Utt(hrs) Audio=10322023.06 | 54.2 | — | — | |
| ASSTGCNPhoneme-based=false, Input=Video+38-Points, Extra Data=false, Total Hours=2232025.07 | 55.7 | — | — | |
| CTC/Attention2019.11 | 63.5 | — | 42.1 | |
| Hyb. CTC/Att.Datasets used (Training)=LRS2, LRW, Total # hours (Training)=3892021.10 | 63.5 | — | — | |
| LIBS2019.11 | 65.29 | 41.91 | 45.53 | |
| LIBSDatasets used (Training)=LRS2, LRS3, Total # hours (Training)=6982021.10 | 65.3 | — | — | |
| Zhao et al.Type=Full-Shot, Labeled Utt(hrs) Video=698, Labeled Utt(hrs) Audio=6982023.06 | 65.3 | — | — | |
| WAS2019.11 | 68.19 | 39.72 | 48.28 | |
| Son Chung et al.Type=Full-Shot, Labeled Utt(hrs) Video=2242023.06 | 70.4 | — | — | |
| P-ASRPhoneme-based=true, Input=117-Points, LLM=NLLB(1.3B), Extra Data=false, Total Hours=2232025.07 | 72.2 | — | — |