Speaker Separation on LRS3 synthetic (test)
15.5SDRVoiceFormer (Ours A+V)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VoiceFormer (Ours A+V)Video conditioning=true, Text conditioning=false2025.01 | 15.5 | 0.934 | 2.62 | |
| VoiceFormer (Ours A+V+T)Video conditioning=true, Text conditioning=true2025.01 | 15.5 | 0.935 | 2.63 | |
| Visual VoiceVideo conditioning=true, Text conditioning=false2025.01 | 11.7 | 0.9 | 2.41 | |
| ConversationVideo conditioning=false, Text conditioning=false2025.01 | 10.15 | 0.865 | 2.08 | |
| AVObjectsVideo conditioning=true, Text conditioning=false2025.01 | 9.72 | 0.851 | 2.02 | |
| Noisy inputVideo conditioning=false, Text conditioning=false2025.01 | 1.3 | 0.697 | 1.3 | |
| DenoiserVideo conditioning=false, Text conditioning=false2025.01 | 1.3 | 0.697 | 1.3 |