Audio-Visual Speech Recognition on LRS3 (test)
0.68WERAVUR-LLM
Evaluation Results
| Method | Links | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AVUR-LLMAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=17592026.03 | 0.68 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UASR-LLM-LUnlabeled data=1759h, Labeled data=1759h, Shared across tasks=true2025.10 | 0.69 | — | — | — | — | — | — | — | — | — | — | — | — | — | 2.4 | |
| MMS-LLaMAAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=17592026.03 | 0.72 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MMS-LlamaSL Input Supported=No, Single Model=No, Training Data (h)=1,759, Learning Paradigm=LLM-Based Models, Leverage pseudo-labels from VoxCeleb2=true2025.08 | 0.72 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UASR-LLM-BUnlabeled data=1759h, Labeled data=433h, Shared across tasks=true2025.10 | 0.73 | — | — | — | — | — | — | — | — | — | — | — | — | — | 2.4 | |
| AVUR-LLMAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=4332026.03 | 0.75 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-AVSRUnlabeled data=1759h, Labeled data=1759h, Shared across tasks=false2025.10 | 0.77 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-AVSRAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=17592026.03 | 0.77 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-AVSRSL Input Supported=No, Single Model=No, Training Data (h)=1,759, Learning Paradigm=LLM-Based Models, Leverage pseudo-labels from VoxCeleb2=true2025.08 | 0.77 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UASR-LLM-B*Unlabeled data=1759h, Labeled data=433h, Shared across tasks=true2025.10 | 0.78 | — | — | — | — | — | — | — | — | — | — | — | — | — | 2.6 | |
| UASR-LLM-LUnlabeled data=1759h, Labeled data=433h, Shared across tasks=true2025.10 | 0.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | 2.7 | |
| UASR-LLM-LUnlabeled data=433h, Labeled data=433h, Shared across tasks=true2025.10 | 0.82 | — | — | — | — | — | — | — | — | — | — | — | — | — | 3.2 | |
| Whisper-FlamingoAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=17592026.03 | 0.86 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Auto-AVSRTraining Hours=34482026.02 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Whisper-FlamingoTraining Hours=3518, fine-tuning=fine-tuned with AV-HuBERT (video) and Whisper (audio)2026.02 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AUTO-AVSRType=A+V, Total Hours=34482023.03 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Auto-AVSRUnlabeled data=-, Labeled data=3.5kh, Shared across tasks=false2025.10 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LP-ConformerUnlabeled data=-, Labeled data=100k, Shared across tasks=false2025.10 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Fast ConformerAudio Encoder=Conformer, Visual Encoder=Conformer, Training Data (h)=16872026.03 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Auto-AVSRAudio Encoder=Conformer, Visual Encoder=Conformer, Training Data (h)=34482026.03 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LP ConformerAudio Encoder=Conformer, Visual Encoder=Conformer, Training Data (h)=100K2026.03 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-MTSKAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=4332026.03 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MMS-LLaMAAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=4332026.03 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Auto-AVSRSL Input Supported=No, Single Model=No, Training Data (h)=3,448, Learning Paradigm=Supervised Learning, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LP ConfSL Input Supported=No, Single Model=No, Training Data (h)=100,000, Learning Paradigm=Supervised Learning, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MMS-LlamaSL Input Supported=No, Single Model=No, Training Data (h)=433, Learning Paradigm=LLM-Based Models, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 0.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UASR-LLM-BUnlabeled data=433h, Labeled data=433h, Shared across tasks=true2025.10 | 0.93 | — | — | — | — | — | — | — | — | — | — | — | — | — | 3.4 | |
| OursSL Input Supported=Yes, Single Model=Yes, Training Data (h)=433, Learning Paradigm=LLM-Based Models, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 0.93 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-AVSRUnlabeled data=1759h, Labeled data=433h, Shared across tasks=false2025.10 | 0.95 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-AVSRAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=4332026.03 | 0.95 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-AVSRSL Input Supported=No, Single Model=No, Training Data (h)=433, Learning Paradigm=LLM-Based Models, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 0.95 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AUTO-AVSRType=A+V, Total Hours=19022023.03 | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Auto-AVSRUnlabeled data=-, Labeled data=1.9kh, Shared across tasks=false2025.10 | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Auto-AVSRAudio Encoder=Conformer, Visual Encoder=Conformer, Training Data (h)=19022026.03 | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Auto-AVSRSL Input Supported=No, Single Model=No, Training Data (h)=1,902, Learning Paradigm=Supervised Learning, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| USR-LUnlabeled data=1759h, Labeled data=433h, Shared across tasks=true2025.10 | 1.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| USRUnlabeled data=1759h, Labeled data=1759h, Shared across tasks=true2025.10 | 1.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Whisper-FlamingoAudio Encoder=Whisper, Visual Encoder=AV-HuBERT, Training Data (h)=4332026.03 | 1.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| USRSL Input Supported=No, Single Model=Yes, Training Data (h)=1,759, Learning Paradigm=Self- or Semi-Supervised Learning, Leverage pseudo-labels from VoxCeleb2=true2025.08 | 1.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VATLM-LUnlabeled data=1759h, Labeled data=433h, Shared across tasks=false2025.10 | 1.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VATLMSL Input Supported=No, Single Model=No, Training Data (h)=433, Learning Paradigm=Self- or Semi-Supervised Learning, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 1.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-data2vec-LUnlabeled data=1759h, Labeled data=433h, Shared across tasks=false2025.10 | 1.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| USR-BUnlabeled data=1759h, Labeled data=433h, Shared across tasks=true2025.10 | 1.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DistillAV-LUnlabeled data=1759h, Labeled data=433h, Shared across tasks=false2025.10 | 1.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | 6.9 | |
| u-HuBERTSL Input Supported=No, Single Model=Yes, Training Data (h)=433, Learning Paradigm=Self- or Semi-Supervised Learning, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 1.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-HuBERTTraining Hours=21922026.02 | 1.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-HUBERTType=A+V, Total Hours=17592023.03 | 1.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-HuBERT-LUnlabeled data=1759h, Labeled data=433h, Shared across tasks=false2025.10 | 1.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-data2vec-BUnlabeled data=1759h, Labeled data=433h, Shared across tasks=false2025.10 | 1.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-HuBERTAudio Encoder=Transformer, Visual Encoder=Transformer, Training Data (h)=4332026.03 | 1.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-HuBERTSL Input Supported=No, Single Model=No, Training Data (h)=433, Learning Paradigm=Self- or Semi-Supervised Learning, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 1.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CMAAudio Encoder=Transformer, Visual Encoder=Transformer, Training Data (h)=4332026.03 | 1.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoBRATraining Hours=6642026.02 | 1.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViT3D-CMType=A+V, Total Hours=900002023.03 | 1.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| USRUnlabeled data=433h, Labeled data=433h, Shared across tasks=true2025.10 | 1.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DistillAV-LUnlabeled data=433h, Labeled data=433h, Shared across tasks=false2025.10 | 1.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | 8.6 | |
| ViT3D-CMAudio Encoder=Conformer, Visual Encoder=Conformer, Training Data (h)=90K2026.03 | 1.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViT3D-CMSL Input Supported=No, Single Model=No, Training Data (h)=90,000, Learning Paradigm=Supervised Learning, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 1.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DistillAV-BUnlabeled data=1759h, Labeled data=433h, Shared across tasks=false2025.10 | 1.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | 7.8 | |
| AV-data2vec-BUnlabeled data=433h, Labeled data=433h, Shared across tasks=false2025.10 | 1.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DistillAV-BUnlabeled data=433h, Labeled data=433h, Shared across tasks=false2025.10 | 1.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | 9.6 | |
| CM-seq2seqTraining Hours=596, mode=baseline2026.02 | 2.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CM-seq2seqType=A+V, Extra Data=No, Total Hours=4382023.03 | 2.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CM-seq2seqAudio Encoder=Conformer, Visual Encoder=-, Training Data (h)=4332026.03 | 2.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-HuBERT-LUnlabeled data=433h, Labeled data=433h, Shared across tasks=false2025.10 | 2.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV2vec-MLMUnlabeled data=433h, Labeled data=433h, Shared across tasks=false2025.10 | 2.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AV-data2vecAudio Encoder=Transformer, Visual Encoder=Transformer, Training Data (h)=4332026.03 | 2.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AVBCspeech enhancement=enabled2026.01 | 3.9 | — | — | 8.5 | 4.5 | 3.2 | 2.6 | 2.4 | 2.1 | — | — | — | — | — | — | |
| AV-RelScore2026.01 | 4.3 | — | — | 9 | 4.8 | 3.3 | 2.9 | 2.9 | 2.8 | — | — | — | — | — | — | |
| Makino et al.Unlabeled data=-, Labeled data=31k, Shared across tasks=false2025.10 | 4.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RNN-TSL Input Supported=No, Single Model=Yes, Training Data (h)=31,000, Learning Paradigm=Supervised Learning, Leverage pseudo-labels from VoxCeleb2=false2025.08 | 4.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RNN-TType=A+V, Total Hours=310002023.03 | 4.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AVBCspeech enhancement=disabled2026.01 | 5.6 | — | — | 12.8 | 6.7 | 4.6 | 3.8 | 3.1 | 2.3 | — | — | — | — | — | — | |
| Joint AVSE-AVSR2026.01 | 6.5 | — | — | 19.4 | 8 | 4.1 | 2.9 | 2.4 | 2 | — | — | — | — | — | — | |
| V-CAFE2026.01 | 8.4 | — | — | 19.3 | 12.5 | 8.4 | 4 | 3 | 2.9 | — | — | — | — | — | — | |
| Conformer2026.01 | 9.6 | — | — | 22.3 | 14.6 | 8.3 | 5.4 | 3.6 | 3.2 | — | — | — | — | — | — | |
| EG-Seq2Seq2026.01 | 9.9 | — | — | 27.3 | 12.2 | 5.8 | 3.9 | 3.3 | 6.8 | — | — | — | — | — | — | |
| ASR + VSR oracle onb / ocpInput=A + V2025.12 | — | — | — | — | — | — | — | — | — | 7.9 | 5.6 | 1.6 | 1.9 | 4.2 | — | |
| ASR oracle onb / ocpInput=A2025.12 | — | — | — | — | — | — | — | — | — | 26.9 | 19.6 | 4.5 | 5 | 14 | — | |
| AV-HuBERT BaseNoise Type=Babble, Params=160M2026.04 | — | — | 41.45 | 23.86 | 10.89 | 5.97 | 5 | — | — | — | — | — | — | — | — | |
| AV-HuBERT BaseNoise Type=Speech, Params=160M2026.04 | — | — | 15.06 | 9.8 | 6.74 | 5.39 | 4.74 | — | — | — | — | — | — | — | — | |
| AV-HuBERT BaseNoise Type=Music, Params=160M2026.04 | — | — | 17.36 | 10.01 | 6.74 | 5.08 | 4.5 | — | — | — | — | — | — | — | — | |
| AV-HuBERT BaseNoise Type=Noise, Params=160M2026.04 | — | — | 15.44 | 9.03 | 6.65 | 5.36 | 4.51 | — | — | — | — | — | — | — | — | |
| AV-HuBERT LargeNoise Type=Babble, Params=477M2026.04 | — | — | 30 | 13.45 | 4.78 | 2.52 | 1.91 | — | — | — | — | — | — | — | — | |
| AV-HuBERT LargeNoise Type=Speech, Params=477M2026.04 | — | — | 13.59 | 4.7 | 3.17 | 1.98 | 1.67 | — | — | — | — | — | — | — | — | |
| AV-HuBERT LargeNoise Type=Music, Params=477M2026.04 | — | — | 10.23 | 4.83 | 2.71 | 1.93 | 1.72 | — | — | — | — | — | — | — | — | |
| AV-HuBERT LargeNoise Type=Noise, Params=477M2026.04 | — | — | 9.82 | 4.88 | 2.55 | 1.87 | 1.72 | — | — | — | — | — | — | — | — | |
| BRAVEn-largeInput=V2025.12 | — | — | — | — | — | — | — | — | — | — | — | — | — | 31.9 | — | |
| Cheng et al.modality=audio-visual2025.03 | — | — | 34.9 | 16.6 | 5.8 | 2.6 | 2 | — | — | — | — | — | — | — | — | |
| DualHypInput=A + V2025.12 | — | — | — | — | — | — | — | — | — | 16.3 | 18.2 | 5.6 | 5.5 | 11.4 | — | |
| DualHyp + RelPromptInput=A + V2025.12 | — | — | — | — | — | — | — | — | — | 14.9 | 16.2 | 5.7 | 5.1 | 10.5 | — | |
| GERInput=A2025.12 | — | — | — | — | — | — | — | — | — | 32.4 | 35.4 | 7.6 | 8 | 20.9 | — | |
| GER w/ Auto-AVSRInput=AV2025.12 | — | — | — | — | — | — | — | — | — | 17.9 | 45.6 | 14.2 | 11 | 22.2 | — | |
| LipGERInput=AV2025.12 | — | — | — | — | — | — | — | — | — | 32.4 | 34.4 | 7.6 | 8.1 | 20.6 | — | |
| MMS-LLaMAmodality=audio-visual2025.03 | — | 17.1 | 13.8 | 7 | 1.9 | 1 | 0.76 | 0.75 | — | — | — | — | — | — | — | |
| RobustGERInput=A2025.12 | — | — | — | — | — | — | — | — | — | 32.5 | 36 | 7.7 | 8.1 | 21.1 | — | |
| VisG AV-HuBERT BaseNoise Type=Babble, Params=162M2026.04 | — | — | 42.55 | 22.6 | 10.18 | 6.03 | 4.65 | — | — | — | — | — | — | — | — | |
| VisG AV-HuBERT BaseNoise Type=Speech, Params=162M2026.04 | — | — | 13.55 | 8.78 | 6.6 | 5.33 | 4.66 | — | — | — | — | — | — | — | — | |
| VisG AV-HuBERT BaseNoise Type=Music, Params=162M2026.04 | — | — | 16.6 | 9.36 | 6.12 | 5 | 4.3 | — | — | — | — | — | — | — | — | |
| VisG AV-HuBERT BaseNoise Type=Noise, Params=162M2026.04 | — | — | 14.72 | 8.87 | 5.59 | 5.27 | 4.66 | — | — | — | — | — | — | — | — |