Visual-only Speech Recognition on LRS2 (test)
12.6WERUSR 2.0
Evaluation Results
| Method | Links | |
|---|---|---|
| USR 2.0Labelled hours=656, Unlabelled hours=2,649, Lang. model=No, Shared params=Yes, Training Protocol=Self/semi-supervised2026.02 | 12.6 | |
| Auto-AVSRLabelled hours=3,448, Language model=true, Shared params=false2024.11 | 14.6 | |
| Auto-AVSRLabelled hours=3,448, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 14.6 | |
| Auto-AVSRPhoneme-based=No, Input=Video, LLM=None, Extra Data=Yes, Training Data (h)=34482026.05 | 14.6 | |
| USRLabelled hours=223, Unlabelled hours=1,759, Language model=true, Shared params=true2024.11 | 15.4 | |
| USRLabelled hours=223, Unlabelled hours=1,759, Lang. model=Yes, Shared params=Yes, Training Protocol=Self/semi-supervised2026.02 | 15.4 | |
| USRLabelled hours=223, Unlabelled hours=1,759, Language model=false, Shared params=true2024.11 | 16 | |
| RAVEn (Large) w/ self-trainingEncoder=Transformer, LM=true, Unlab hours=1,759, Lab hours=223, Supervision regime=self-supervised2022.12 | 17.9 | |
| RAVEn w/ STLabelled hours=223, Unlabelled hours=1,759, Language model=true, Shared params=false2024.11 | 17.9 | |
| RAVEn w/ STLabelled hours=223, Unlabelled hours=1,759, Lang. model=Yes, Shared params=No, Training Protocol=Self/semi-supervised2026.02 | 17.9 | |
| RAVEn (Large) w/ self-trainingEncoder=Transformer, LM=false, Unlab hours=1,759, Lab hours=223, Supervision regime=self-supervised2022.12 | 19.3 | |
| VALLRPhoneme-based=Yes, Input=Video, LLM=LLaMA, Extra Data=No, Training Data (h)=282026.05 | 20.8 | |
| Prajwal et al. (2022)Encoder=Transformer, LM=true, Lab hours=2,676*, Supervision regime=supervised2022.12 | 22.6 | |
| VTPLabelled hours=2,676, Language model=true, Shared params=false2024.11 | 22.6 | |
| VTPLabelled hours=2,676, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 22.6 | |
| VTPPhoneme-based=No, Input=Video, LLM=None, Extra Data=Yes, Training Data (h)=26762026.05 | 22.6 | |
| RAVEn (Large)Encoder=Transformer, LM=false, Unlab hours=1,759, Lab hours=223, Supervision regime=self-supervised2022.12 | 23.2 | |
| RAVEnLabelled hours=223, Unlabelled hours=1,759, Language model=false, Shared params=false2024.11 | 23.2 | |
| RAVEnLabelled hours=223, Unlabelled hours=1,759, Lang. model=No, Shared params=No, Training Protocol=Self/semi-supervised2026.02 | 23.2 | |
| HP-VSR-ResFiLMPhoneme-based=Yes, Input=Video+Head Pose, LLM=NLLB(1.3B), Extra Data=No, Training Data (h)=2232026.05 | 25 | |
| Ma et al. (2022)Encoder=Conformer, LM=true, Unlab hours=641, Lab hours=818, Supervision regime=semi-supervised2022.12 | 25.5 | |
| CM-auxLabelled hours=1,459, Language model=true, Shared params=false2024.11 | 25.5 | |
| CM-auxLabelled hours=1,459, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 25.5 | |
| HP-VSR-FiLMFuse(L3–4)Phoneme-based=Yes, Input=Video+Head Pose, LLM=NLLB(1.3B), Extra Data=No, Training Data (h)=2232026.05 | 25.7 | |
| PV-ASRPhoneme-based=Yes, Input=Video+117-Points, LLM=NLLB(1.3B), Extra Data=No, Training Data (h)=2232026.05 | 26.9 | |
| Ma et al. (2022)Encoder=Conformer, LM=true, Lab hours=818, Supervision regime=supervised2022.12 | 27.3 | |
| Auto-AVSRLabelled hours=818, Language model=true, Shared params=false2024.11 | 27.9 | |
| Auto-AVSRLabelled hours=818, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 27.9 | |
| Auto-AVSRPhoneme-based=No, Input=Video, LLM=None, Extra Data=Yes, Training Data (h)=8182026.05 | 27.9 | |
| Prajwal et al. (2022)Encoder=Transformer, LM=true, Lab hours=698, Supervision regime=supervised2022.12 | 28.9 | |
| VTPLabelled hours=698, Language model=true, Shared params=false2024.11 | 28.9 | |
| VTPLabelled hours=698, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 28.9 | |
| RAVEn (Base)Encoder=Transformer, LM=false, Unlab hours=433, Lab hours=223, Supervision regime=self-supervised2022.12 | 32.1 | |
| CM-auxPhoneme-based=No, Input=Video, LLM=None, Extra Data=No, Training Data (h)=2232026.05 | 32.9 | |
| HP-VSR-BasePhoneme-based=Yes, Input=Video+Head Pose, LLM=NLLB(1.3B), Extra Data=No, Training Data (h)=2232026.05 | 32.9 | |
| Auto-AVSR Finetuned BaselinePhoneme-based=No, Input=Video, LLM=None, Extra Data=No, Training Data (h)=2232026.05 | 33 | |
| V-ASRPhoneme-based=Yes, Input=Video, LLM=NLLB(1.3B), Extra Data=No, Training Data (h)=2232026.05 | 35.2 | |
| E2E ConformerLanguage Model=Powerful Transformer2022.02 | 37.9 | |
| Ma et al. (2021b)Encoder=Conformer, LM=true, Lab hours=380, Supervision regime=supervised2022.12 | 37.9 | |
| CM-seq2seqLabelled hours=380, Language model=true, Shared params=false2024.11 | 37.9 | |
| CM-seq2seqLabelled hours=380, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 37.9 | |
| Hyb.-Conf.Phoneme-based=No, Input=Video, LLM=None, Extra Data=Yes, Training Data (h)=3812026.05 | 37.9 | |
| Ma et al. (2021a)Encoder=Conformer, LM=true, Unlab hours=433, Lab hours=223, Supervision regime=self-supervised2022.12 | 38.8 | |
| LIRALabelled hours=223, Unlabelled hours=433, Language model=true2024.11 | 38.8 | |
| LiRALabelled hours=223, Unlabelled hours=433, Lang. model=Yes, Shared params=No, Training Protocol=Self/semi-supervised2026.02 | 38.8 | |
| E2E ConformerLanguage Model=External2022.02 | 42.4 | |
| Our ModelLanguage Model=None2022.02 | 43.2 | |
| Pan et al. (2022)Encoder=Transformer, LM=false, Unlab hours=60,000, Lab hours=223, Supervision regime=self-supervised2022.12 | 43.2 | |
| Uni-AVSRLabelled hours=223, Unlabelled hours=60,000, Language model=false, Shared params=false2024.11 | 43.2 | |
| Uni-AVSRLabelled hours=223, Unlabelled hours=60,000, Lang. model=No, Shared params=No, Training Protocol=Self/semi-supervised2026.02 | 43.2 | |
| TM-seq2seqTrained on Vox.=X, Trained on LRS2=GT, Trained on LRS3=GT2019.11 | 48.3 | |
| LF-MMI TDNNModality=Visual-only2020.01 | 48.86 | |
| LF-MMI TDNNLanguage Model=External2022.02 | 48.9 | |
| Yu et al. (2020)Encoder=CNN, LM=true, Lab hours=223, Supervision regime=supervised2022.12 | 48.9 | |
| VTPPhoneme-based=No, Input=Video, LLM=None, Extra Data=Yes, Training Data (h)=6982026.05 | 48.9 | |
| KD-TM2022.02 | 49.2 | |
| Ren et al. (2021)Encoder=Transformer, LM=false, Lab hours=818, Supervision regime=supervised2022.12 | 49.2 | |
| TM-Seq2seqModality=Visual-only2020.01 | 49.8 | |
| TM-seq2seqLanguage Model=External2022.02 | 50 | |
| CTC + KDTrained on Vox.=ASR, Trained on LRS2=GT, Trained on LRS3=ASR2019.11 | 51.3 | |
| Afouras et al. (2020)Encoder=CNN, LM=true, Unlab hours=777, Lab hours=223, Supervision regime=semi-supervised2022.12 | 51.3 | |
| Conv-seq2seqTrained on Vox.=X, Trained on LRS2=GT, Trained on LRS3=GT2019.11 | 51.7 | |
| Conv-seq2seq2022.02 | 51.7 | |
| CTC + KDTrained on Vox.=ASR, Trained on LRS2=ASR/GT, Trained on LRS3=ASR2019.11 | 52.2 | |
| CTC + KDTrained on Vox.=ASR, Trained on LRS2=ASR, Trained on LRS3=ASR2019.11 | 54.2 | |
| TM-CTCLanguage Model=External2022.02 | 54.7 | |
| CE TDNNModality=Visual-only2020.01 | 55.02 | |
| ASSTGCNPhoneme-based=No, Input=Video+38-Points, LLM=None, Extra Data=No, Training Data (h)=2232026.05 | 55.7 | |
| CTC + KDTrained on Vox.=X, Trained on LRS2=ASR/GT, Trained on LRS3=X2019.11 | 57.9 | |
| CTC + KDTrained on Vox.=X, Trained on LRS2=ASR, Trained on LRS3=X2019.11 | 58.2 | |
| CTCTrained on Vox.=X, Trained on LRS2=GT, Trained on LRS3=X2019.11 | 58.5 | |
| CTC/AttentionModality=Visual-only2020.01 | 63.5 | |
| Hyb. CTC/Att.Trained on Vox.=X, Trained on LRS2=GT, Trained on LRS3=X2019.11 | 63.5 | |
| Petridis et al. (2018)Encoder=LSTM, LM=true, Lab hours=380, Supervision regime=supervised2022.12 | 63.5 | |
| TM-CTCModality=Visual-only2020.01 | 65 | |
| LIBS2022.02 | 65.3 | |
| Chung et al. (2017)Encoder=LSTM, LM=true, Lab hours=223, Supervision regime=supervised2022.12 | 70.4 |