Visual Speech Recognition on CMLR
8Best CERVSR model with prediction-based auxiliary tasks
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3+AVSpeech, Training Set=CMLR, Training Sets Total Size (hours)=15202022.02 | 8 | 8.1 | 0.05 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=LRW+LRS2+LRS3, Training Set=CMLR, Training Sets Total Size (hours)=8792022.02 | 8.1 | 8.2 | 0.06 | |
| VSR model with prediction-based auxiliary tasksPre-training Set=None, Training Set=CMLR, Training Sets Total Size (hours)=612022.02 | 9.1 | 9.1 | 0.05 | |
| CTCHPre-training Set=None, Training Set=CMLR, Training Sets Total Size (hours)=612022.02 | 22 | — | — | |
| LIBSPre-training Set=None, Training Set=CMLR, Training Sets Total Size (hours)=612022.02 | 31.3 | — | — | |
| CSSMCMPre-training Set=None, Training Set=CMLR, Training Sets Total Size (hours)=612022.02 | 32.5 | — | — | |
| LipCH-NetPre-training Set=None, Training Set=CMLR, Training Sets Total Size (hours)=612022.02 | 34 | — | — |