Audio-Visual Speech Recognition on LRS3
0.008WERLlama-AVSR
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-AVSRLabelled hours=680,000, Lang. model=true, Shared params=false, Encoder=Whisper2026.02 | 0.008 | |
| Whisper-FlamingoLabelled hours=680,000, Lang. model=true, Shared params=false, Encoder=Whisper2026.02 | 0.008 | |
| USR 2.0Labelled hours=656, Unlabelled hours=2,649, Lang. model=false, Shared params=true2026.02 | 0.008 | |
| Auto-AVSRLabelled hours=3,448, Lang. model=true, Shared params=false2026.02 | 0.009 | |
| LP ConfLabelled hours=100,000, Lang. model=false, Shared params=false2026.02 | 0.009 | |
| Auto-AVSRLabelled hours=1,902, Lang. model=true, Shared params=false2026.02 | 0.01 | |
| USRLabelled hours=433, Unlabelled hours=1,326, Lang. model=true, Shared params=true2026.02 | 0.011 | |
| ViT3D-CMLabelled hours=90,000, Lang. model=false, Shared params=false2026.02 | 0.016 | |
| RNN-TLabelled hours=31,000, Lang. model=false, Shared params=true2026.02 | 0.045 | |
| Whisper-flamingoPreprocessing=Lip crop & align2026.03 | 76 | |
| HumanOmni-SpeakerPreprocessing=Raw video2026.03 | 76 | |
| Llama-AVSRPreprocessing=Lip crop & align2026.03 | 77 | |
| AutoAVSRPreprocessing=Lip crop & align2026.03 | 90 | |
| Llama-SMoPPreprocessing=Lip crop & align2026.03 | 96 |