Visual Speech Recognition on LRS2
14.6Mean WERAutoAVSR
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AutoAVSRVideo Hours=3,448, LM=true, Model Size=base-size2024.06 | 14.6 | — | — | |
| AUTO-AVSRType=V, Total Hours=34482023.03 | 14.6 | — | — | |
| SyncVSRVideo Hours=1,992, LM=true, Model Size=base-size2024.06 | 16.5 | — | — | |
| RAVEn-LargeVideo Hours=1,992/1,759, LM=true, Model Size=large-size2024.06 | 17.9 | — | — | |
| SyncVSRVideo Hours=1,992, LM=false, Model Size=base-size2024.06 | 18.5 | — | — | |
| RAVEn-LargeVideo Hours=1,992/1,759, LM=false, Model Size=large-size2024.06 | 19.3 | — | — | |
| SyncVSRVideo Hours=661, LM=true, Model Size=base-size2024.06 | 20 | — | — | |
| SyncVSRVideo Hours=661, LM=false, Model Size=base-size2024.06 | 22 | — | — | |
| VTPVideo Hours=2,676, LM=true, Model Size=base-size2024.06 | 22.6 | — | — | |
| VTPType=V, Total Hours=26762023.03 | 22.6 | — | — | |
| LMDecoderVideo Hours=1,992, LM=false, Model Size=base-size2024.06 | 23.8 | — | — | |
| VATLM-LargeVideo Hours=1,992/1,759, LM=false, Model Size=large-size2024.06 | 24.3 | — | — | |
| CM-AuxVideo Hours=1,459, LM=true, Model Size=base-size2024.06 | 25.5 | — | — | |
| AV-HUBERT-LargeVideo Hours=1,992/1,759, LM=false, Model Size=large-size2024.06 | 25.5 | — | — | |
| CM-auxType=V, Total Hours=14592023.03 | 25.5 | — | — | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS3+AVSpeech, Training Set=LRS2, Training Sets Total Size (hours)=14592022.02 | 25.8 | 0.4 | 25.5 | |
| CM-AuxVideo Hours=818, LM=true, Model Size=base-size2024.06 | 27.3 | — | — | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS3, Training Set=LRS2, Training Sets Total Size (hours)=8182022.02 | 27.6 | 0.2 | 27.3 | |
| AutoAVSRVideo Hours=818, LM=true, Model Size=base-size2024.06 | 27.9 | — | — | |
| AUTO-AVSRType=V, Total Hours=8182023.03 | 27.9 | — | — | |
| AutoAVSRPreprocessing=Lip crop & align2026.03 | 27.9 | — | — | |
| SyncVSRVideo Hours=223/438, LM=true, Model Size=base-size2024.06 | 28.9 | — | — | |
| VTPVideo Hours=698, LM=true, Model Size=base-size2024.06 | 28.9 | — | — | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW, Training Set=LRS2, Training Sets Total Size (hours)=3802022.02 | 29.5 | 0.4 | 28.7 | |
| HumanOmni-SpeakerPreprocessing=Raw video2026.03 | 29.8 | — | — | |
| VATLMVideo Hours=1,992/1,759, LM=false, Model Size=base-size2024.06 | 30.6 | — | — | |
| SyncVSRVideo Hours=223/438, LM=false, Model Size=base-size2024.06 | 30.7 | — | — | |
| AV-HUBERTVideo Hours=1,992/1,759, LM=false, Model Size=base-size2024.06 | 31.2 | — | — | |
| AV-HuBERT BasePreprocessing=Lip crop & align2026.03 | 31.2 | — | — | |
| RAVEnVideo Hours=661, LM=true, Model Size=base-size2024.06 | 32.1 | — | — | |
| CM-AuxVideo Hours=223/438, LM=true, Model Size=base-size2024.06 | 32.9 | — | — | |
| CM-auxType=V2023.03 | 32.9 | — | — | |
| VSR model with prediction-based auxiliary tasksTraining Set=LRS2, Training Sets Total Size (hours)=2232022.02 | 33.6 | 0.5 | 32.9 | |
| LIRAVideo Hours=661, LM=true, Model Size=base-size2024.06 | 38.8 | — | — | |
| CM-Seq2SeqVideo Hours=223/438, LM=true, Model Size=base-size2024.06 | 39.1 | — | — | |
| CM-seq2seqType=V, Extra Data=false, Total Hours=2232023.03 | 39.1 | — | — | |
| CTC/AttentionType=V, Total Hours=600002023.03 | 43.2 | — | — | |
| MVMVideo Hours=818, LM=false, Model Size=base-size2024.06 | 44.5 | — | — | |
| TM-Seq2SeqVideo Hours=1,391, LM=true, Model Size=base-size2024.06 | 48.3 | — | — | |
| TM-seq2seqType=V, Total Hours=13912023.03 | 48.3 | — | — | |
| TDNNVideo Hours=223, LM=true, Model Size=base-size2024.06 | 48.9 | — | — | |
| TDNNType=V2023.03 | 48.9 | — | — | |
| KD-Seq2SeqVideo Hours=818, LM=false, Model Size=base-size2024.06 | 49.2 | — | — | |
| KD-seq2seqType=V, Total Hours=8182023.03 | 49.2 | — | — | |
| KD + CTCVideo Hours=995, LM=true, Model Size=base-size2024.06 | 51.3 | — | — | |
| KD + CTCType=V, Total Hours=9952023.03 | 51.3 | — | — | |
| CTC/AttentionType=V, Total Hours=3802023.03 | 63.5 | — | — | |
| CTC/AttentionPreprocessing=Lip crop & align2026.03 | 63.5 | — | — | |
| MV-WASType=V2023.03 | 70.4 | — | — | |
| CM-seq2seqPre-training Set=LRW, Training Set=LRS2, Training Sets Total Size (hours)=3802022.02 | — | — | 37.9 | |
| CTC/Att.Pre-training Set=LRW, Training Set=LRS2, Training Sets Total Size (hours)=3802022.02 | — | — | 63.5 | |
| KD + CTCPre-training Set=VoxCeleb2clean+LRS3, Training Set=LRS2, Training Sets Total Size (hours)=9952022.02 | — | — | 51.3 | |
| KD-seq2seqPre-training Set=LRW+LRS3, Training Set=LRS2, Training Sets Total Size (hours)=8182022.02 | — | — | 49.2 | |
| MV-WASTraining Set=LRS2, Training Sets Total Size (hours)=2232022.02 | — | — | 70.4 | |
| TDNNTraining Set=LRS2, Training Sets Total Size (hours)=2232022.02 | — | — | 48.9 | |
| TM-seq2seqPre-training Set=MVLRS-LRS3, Training Set=LRS2, Training Sets Total Size (hours)=13912022.02 | — | — | 48.3 |