Audio-Visual Speech Recognition on LRS2
8.9Error Rate (B)CAV2vec
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CAV2vecVisual corruption type=object occlusion + noise2025.12 | 8.9 | 4.4 | 5.1 | 4.9 | 2.7 | — | |
| AV-RelScoreVisual corruption type=object occlusion + noise2025.12 | 11.1 | 4.8 | 5.9 | 5.5 | 2.9 | — | |
| AV-data2vecVisual corruption type=object occlusion + noise2025.12 | 11.5 | 5.6 | 6.5 | 6.2 | 3 | — | |
| AV-HuBERTVisual corruption type=object occlusion + noise2025.12 | 11.6 | 5.3 | 6.1 | 6 | 3 | — | |
| AutoAVSRPreprocessing=Lip crop & align2026.03 | — | — | — | — | — | 1.5 | |
| CTC/AttentionPreprocessing=Lip crop & align2026.03 | — | — | — | — | — | 7 | |
| HumanOmni-SpeakerPreprocessing=Raw video2026.03 | — | — | — | — | — | 1.36 | |
| Whisper-flamingoPreprocessing=Lip crop & align2026.03 | — | — | — | — | — | 1.4 |