Visual Speech Recognition on LRS3 High-Resource, 433h labelled v1 (test)
0.009WERChang et al.
Evaluation Results
| Method | Links | |
|---|---|---|
| Chang et al.Label Data=100kh, Criterion=Transducer2025.02 | 0.009 | |
| USRPre-train data=LRS3+Vox2, Shared params=true, Model scale=Large2024.11 | 0.011 | |
| VATLMPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.012 | |
| VATLMUnlabeled Data=1759h, Label Data=433h, Encoder Size=325M, Criterion=CE2025.02 | 0.012 | |
| u-HuBERTLabel Data=433h, Encoder Size=325M, Criterion=CE2025.02 | 0.012 | |
| USRPre-train data=LRS3+Vox2, Shared params=true, Model scale=Base(+)2024.11 | 0.013 | |
| AV-data2vecPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.013 | |
| u-HuBERTPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.013 | |
| AV-data2vecEncoder Size=325M, Criterion=CE, Setting=Large Block 22025.02 | 0.013 | |
| DistillAVEncoder Size=325M, Criterion=CE+KD, Setting=Large Block 22025.02 | 0.013 | |
| AV-data2vecPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.014 | |
| AV-HuBERTPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.014 | |
| AV-data2vecEncoder Size=103M, Criterion=CE, Setting=Base Block 22025.02 | 0.014 | |
| AV-HuBERTEncoder Size=325M, Criterion=CE, Setting=Large Block 22025.02 | 0.014 | |
| USRPre-train data=LRS3, Shared params=true, Model scale=Base(+)2024.11 | 0.016 | |
| Serdyuk et al.Label Data=90kh, Criterion=Transducer2025.02 | 0.016 | |
| DistillAVEncoder Size=325M, Criterion=CE+KD2025.02 | 0.016 | |
| VATLMPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.017 | |
| VATLMUnlabeled Data=1759h, Label Data=433h, Encoder Size=103M, Criterion=CE2025.02 | 0.017 | |
| DistillAVEncoder Size=103M, Criterion=CE+KD, Setting=Base Block 22025.02 | 0.017 | |
| AV-data2vecUnlabeled Data=433h, Label Data=433h, Encoder Size=325M, Criterion=CE2025.02 | 0.017 | |
| AV-data2vecPre-train data=LRS3, Shared params=false, Model scale=Base(+)2024.11 | 0.018 | |
| AV-HuBERTPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.018 | |
| AV-data2vecUnlabeled Data=433h, Label Data=433h, Encoder Size=103M, Criterion=CE2025.02 | 0.018 | |
| DistillAVEncoder Size=103M, Criterion=CE+KD2025.02 | 0.018 | |
| AV-HuBERTEncoder Size=103M, Criterion=CE, Setting=Base Block 22025.02 | 0.018 | |
| AV2vec-MLMEncoder Size=103M, Criterion=CE2025.02 | 0.025 | |
| AV-HuBERTEncoder Size=325M, Criterion=CE2025.02 | 0.025 | |
| AV-HuBERTPre-train data=LRS3, Shared params=false, Model scale=Base(+)2024.11 | 0.028 | |
| AV-HuBERTEncoder Size=103M, Criterion=CE2025.02 | 0.028 | |
| Makino et al.Label Data=31kh, Criterion=Transducer2025.02 | 0.045 | |
| Chang et al.Label Data=100kh, Criterion=Transducer2025.02 | 0.128 | |
| Serdyuk et al.Label Data=90kh, Criterion=Transducer2025.02 | 0.17 | |
| USRPre-train data=LRS3+Vox2, Shared params=true, Model scale=Large2024.11 | 0.223 | |
| Lip2VecPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.26 | |
| Lip2VecEncoder Size=727M, Criterion=CTC+KD, Setting=Self-supervised model (Large) Block 22025.02 | 0.26 | |
| DistillAVEncoder Size=325M, Criterion=CE+KD, Setting=Self-supervised model (Large) Block 22025.02 | 0.262 | |
| USRPre-train data=LRS3+Vox2, Shared params=true, Model scale=Base(+)2024.11 | 0.265 | |
| BRAVEnPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.266 | |
| u-HuBERTLabel Data=433h, Encoder Size=325M, Criterion=CE2025.02 | 0.272 | |
| RAVEnPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.282 | |
| RAVenEncoder Size=671M, Criterion=CTC+CE2025.02 | 0.282 | |
| VATLMPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.284 | |
| VATLMUnlabeled Data=1759h, Label Data=433h, Encoder Size=325M, Criterion=CE2025.02 | 0.284 | |
| AV-data2vecPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.285 | |
| AV-data2vecEncoder Size=325M, Criterion=CE, Setting=Self-supervised model (Large) Block 22025.02 | 0.285 | |
| AV-HuBERTPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.286 | |
| AV-HuBERTEncoder Size=325M, Criterion=CE, Setting=Self-supervised model (Large) Block 22025.02 | 0.286 | |
| BRAVEnPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.288 | |
| u-HuBERTPre-train data=LRS3+Vox2, Shared params=false, Model scale=Large2024.11 | 0.291 | |
| Prajwal et al.Label Data=2726h, Criterion=CE2025.02 | 0.307 | |
| DistillAVEncoder Size=103M, Criterion=CE+KD, Setting=Self-supervised model (Base) Block 22025.02 | 0.314 | |
| DistillAVEncoder Size=325M, Criterion=CE+KD, Setting=Self-supervised model (Large) Block 12025.02 | 0.315 | |
| AV-data2vecEncoder Size=103M, Criterion=CE, Setting=Self-supervised model (Base) Block 22025.02 | 0.327 | |
| AV-data2vecPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.329 | |
| RAVEnPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.331 | |
| RAVenEncoder Size=97M, Criterion=CTC+CE, Setting=Self-supervised model (Base) Block 22025.02 | 0.331 | |
| Makino et al.Label Data=31kh, Criterion=Transducer2025.02 | 0.336 | |
| Lip2VecPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.341 | |
| Lip2VecEncoder Size=240M, Criterion=CTC+KD, Setting=Self-supervised model (Base) Block 22025.02 | 0.341 | |
| VATLMPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.342 | |
| VATLMUnlabeled Data=1759h, Label Data=433h, Encoder Size=103M, Criterion=CE2025.02 | 0.342 | |
| USRPre-train data=LRS3, Shared params=true, Model scale=Base(+)2024.11 | 0.343 | |
| DistillAVEncoder Size=103M, Criterion=CE+KD, Setting=Self-supervised model (Base) Block 12025.02 | 0.343 | |
| AV2vec-MLMEncoder Size=103M, Criterion=CE2025.02 | 0.344 | |
| AV-HuBERTPre-train data=LRS3+Vox2, Shared params=false, Model scale=Base(+)2024.11 | 0.348 | |
| AV-HuBERTEncoder Size=103M, Criterion=CE, Setting=Self-supervised model (Base) Block 22025.02 | 0.348 | |
| BRAVEnPre-train data=LRS3, Shared params=false, Model scale=Base(+)2024.11 | 0.36 | |
| AV-data2vecUnlabeled Data=433h, Label Data=433h, Encoder Size=325M, Criterion=CE2025.02 | 0.374 | |
| AV-data2vecPre-train data=LRS3, Shared params=false, Model scale=Base(+)2024.11 | 0.39 | |
| AV-data2vecUnlabeled Data=433h, Label Data=433h, Encoder Size=103M, Criterion=CE2025.02 | 0.39 | |
| RAVEnPre-train data=LRS3, Shared params=false, Model scale=Base(+)2024.11 | 0.391 | |
| RAVenEncoder Size=97M, Criterion=CTC+CE2025.02 | 0.391 | |
| AV-HuBERTEncoder Size=325M, Criterion=CE2025.02 | 0.416 | |
| Lip2VecPre-train data=LRS3, Shared params=false, Model scale=Base(+)2024.11 | 0.42 | |
| Lip2VecEncoder Size=240M, Criterion=CTC+KD2025.02 | 0.42 | |
| AV-HuBERTPre-train data=LRS3, Shared params=false, Model scale=Base(+)2024.11 | 0.44 | |
| AV-HuBERTEncoder Size=103M, Criterion=CE2025.02 | 0.44 | |
| Lip2VecEncoder Size=727M, Criterion=CTC+KD2025.02 | 0.501 | |
| Afouras et al.Label Data=1519h, Criterion=CE2025.02 | 0.589 |