Speech Reconstruction on ReSSInt laryngeal subjects (test)
40.5WERMasked Multimodal Speech Synthesis Framework
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Masked Multimodal Speech Synthesis FrameworkModality=sEMG + Lips (Multimodal), Setting=Full Model2026.06 | 40.5 | 76.3 | 56.6 | |
| Masked Multimodal Speech Synthesis FrameworkModality=sEMG + Lips (Multimodal), Setting=w/o Random Masking2026.06 | 44 | 76.2 | 57.2 | |
| Masked Multimodal Speech Synthesis FrameworkModality=Lip-only, Setting=w/o Random Masking2026.06 | 54.5 | 72.8 | 56.8 | |
| Masked Multimodal Speech Synthesis FrameworkModality=Lip-only, Setting=Full model2026.06 | 57.2 | 71.4 | 54.8 | |
| Masked Multimodal Speech Synthesis FrameworkModality=sEMG + Lips (Multimodal), Setting=w/o Phone Loss2026.06 | 58.8 | — | 55.5 | |
| Masked Multimodal Speech Synthesis FrameworkModality=Lip-only, Setting=w/o Phone Loss2026.06 | 62 | — | 56.1 | |
| Masked Multimodal Speech Synthesis FrameworkModality=sEMG-only, Setting=w/o Random Masking2026.06 | 83.2 | 63.6 | 52 | |
| Masked Multimodal Speech Synthesis FrameworkModality=sEMG-only, Setting=w/o Phone Loss2026.06 | 93.6 | — | 51.1 | |
| Masked Multimodal Speech Synthesis FrameworkModality=sEMG-only, Setting=Full model2026.06 | 94.1 | 58.4 | 48.3 |