Speech Recognition on LRS2 (test)
1.3WERWhisper-Flamingo
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Whisper-FlamingoLabelled hours=680,000, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 1.3 | — | — | |
| USR 2.0Labelled hours=656, Unlabelled hours=2,649, Lang. model=No, Shared params=Yes, Training Protocol=Self/semi-supervised2026.02 | 1.3 | — | — | |
| Auto-AVSRLabelled hours=3,448, Language model=true, Shared params=false2024.11 | 1.5 | — | — | |
| Auto-AVSRLabelled hours=3,448, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 1.5 | — | — | |
| USRLabelled hours=223, Unlabelled hours=1,759, Language model=true, Shared params=true2024.11 | 1.9 | — | — | |
| USRLabelled hours=223, Unlabelled hours=1,759, Lang. model=Yes, Shared params=Yes, Training Protocol=Self/semi-supervised2026.02 | 1.9 | — | — | |
| USRLabelled hours=223, Unlabelled hours=1,759, Language model=false, Shared params=true2024.11 | 2 | — | — | |
| RAVEn (Large) w/ self-trainingEncoder=Transformer, LM=false, Unlab hours=1,759, Lab hours=223, Supervision regime=self-supervised2022.12 | 2.3 | — | — | |
| RAVEn (Large) w/ self-trainingEncoder=Transformer, LM=true, Unlab hours=1,759, Lab hours=223, Supervision regime=self-supervised2022.12 | 2.3 | — | — | |
| RAVEn w/ STLabelled hours=223, Unlabelled hours=1,759, Language model=true, Shared params=false2024.11 | 2.3 | — | — | |
| RAVEn w/ STLabelled hours=223, Unlabelled hours=1,759, Lang. model=Yes, Shared params=No, Training Protocol=Self/semi-supervised2026.02 | 2.3 | — | — | |
| RAVEn (Large)Encoder=Transformer, LM=false, Unlab hours=1,759, Lab hours=223, Supervision regime=self-supervised2022.12 | 2.5 | — | — | |
| RAVEnLabelled hours=223, Unlabelled hours=1,759, Language model=false, Shared params=false2024.11 | 2.5 | — | — | |
| RAVEnLabelled hours=223, Unlabelled hours=1,759, Lang. model=No, Shared params=No, Training Protocol=Self/semi-supervised2026.02 | 2.5 | — | — | |
| Auto-AVSRLabelled hours=818, Language model=true, Shared params=false2024.11 | 2.6 | — | — | |
| Auto-AVSRLabelled hours=818, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 2.6 | — | — | |
| Our ModelLanguage Model=None2022.02 | 2.7 | — | — | |
| Pan et al. (2022)Encoder=Transformer, LM=false, Unlab hours=60,000, Lab hours=223, Supervision regime=self-supervised2022.12 | 2.7 | — | — | |
| Uni-AVSRLabelled hours=223, Unlabelled hours=60,000, Language model=false, Shared params=false2024.11 | 2.7 | — | — | |
| Uni-AVSRLabelled hours=223, Unlabelled hours=60,000, Lang. model=No, Shared params=No, Training Protocol=Self/semi-supervised2026.02 | 2.7 | — | — | |
| E2E ConformerLanguage Model=Powerful Transformer2022.02 | 3.9 | — | — | |
| Ma et al. (2021b)Encoder=Conformer, LM=true, Lab hours=380, Supervision regime=supervised2022.12 | 3.9 | — | — | |
| RAVEn (Base)Encoder=Transformer, LM=false, Unlab hours=433, Lab hours=223, Supervision regime=self-supervised2022.12 | 3.9 | — | — | |
| CM-seq2seqLabelled hours=380, Language model=true, Shared params=false2024.11 | 3.9 | — | — | |
| CM-seq2seqLabelled hours=380, Unlabelled hours=-, Lang. model=Yes, Shared params=No, Training Protocol=Supervised2026.02 | 3.9 | — | — | |
| LF-MMI TDNNLanguage Model=External2022.02 | 6.7 | — | — | |
| Yu et al. (2020)Encoder=CNN, LM=true, Lab hours=223, Supervision regime=supervised2022.12 | 6.7 | — | — | |
| LF-MMI TDNNModality=Audio-only2020.01 | 6.71 | — | — | |
| hybrid CTC/attention architectureStream=Audio + Video (A + V), Fusion strategy=Early Fusion, Visual frame rate=50 fps, Audio frame rate=50 fps2018.09 | 7 | 3.6 | — | |
| ClozeGERBackbone=LLaMA-2-7b, Source Speech=false, LoRA=true, Logits Calibration=true, Post-processing=true2024.05 | 7.4 | — | 3,980 | |
| ClozeGERBackbone=SpeechGPT, Source Speech=true, LoRA=true, Logits Calibration=true, Post-processing=true2024.05 | 7.6 | — | 3,820 | |
| Afouras et al.Stream=Audio + Video (A + V), Pre-training=non-publicly available dataset2018.09 | 8.2 | — | — | |
| CTC/attentionLanguage Model=External2022.02 | 8.2 | — | — | |
| hybrid CTC/attention architectureStream=Audio-only (A), Audio frame rate=100 fps2018.09 | 8.3 | 4.4 | — | |
| CTC/AttentionModality=Audio-only2020.01 | 8.3 | — | — | |
| Petridis et al. (2018)Encoder=LSTM, LM=true, Lab hours=380, Supervision regime=supervised2022.12 | 8.3 | — | — | |
| hybrid CTC/attention architectureStream=Audio + Video (A + V), Fusion strategy=Late Fusion, Visual frame rate=50 fps, Audio frame rate=50 fps2018.09 | 8.5 | 4.7 | — | |
| Afouras et al.Stream=Audio-only (A), Pre-training=non-publicly available dataset2018.09 | 9.7 | — | — | |
| TM-seq2seqLanguage Model=External2022.02 | 9.7 | — | — | |
| TM-CTCLanguage Model=External2022.02 | 10.1 | — | — | |
| CE TDNNModality=Audio-only2020.01 | 10.17 | — | — | |
| TM-Seq2seqModality=Audio-only2020.01 | 10.5 | — | — | |
| Whisper Baseline2024.05 | 12.3 | — | — | |
| TM-CTCModality=Audio-only2020.01 | 15.3 | — | — | |
| Method [30]Stream=Audio-only (A)2018.09 | 29.9 | 14.3 | — | |
| Method [30]Stream=Audio + Video (A + V)2018.09 | 30.5 | 14.1 | — | |
| Afouras et al.Stream=Video-only (V), Pre-training=non-publicly available dataset2018.09 | 50 | — | — | |
| hybrid CTC/attention architectureStream=Video-only (V), Visual frame rate=50 fps2018.09 | 63.5 | 42.1 | — | |
| Method [15]Stream=Video-only (V)2018.09 | 70.4 | — | — |