Speaker Separation on LRS2 synthetic (test)
14.2SDRVoiceFormer (Ours A+V+T)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VoiceFormer (Ours A+V+T)Video conditioning=true, Text conditioning=true2025.01 | 14.2 | 91.7 | 2.41 | |
| VoiceFormer (Ours A+V)Video conditioning=true, Text conditioning=false2025.01 | 14.1 | 91.3 | 2.36 | |
| Visual VoiceVideo conditioning=true, Text conditioning=false2025.01 | 10.8 | 88.4 | 2.16 | |
| ConversationVideo conditioning=false, Text conditioning=false2025.01 | 9.25 | 84 | 1.91 | |
| AVObjectsVideo conditioning=true, Text conditioning=false2025.01 | 8.86 | 83.9 | 1.94 | |
| Noisy inputVideo conditioning=false, Text conditioning=false2025.01 | 0.2 | 66.6 | 1.17 | |
| DenoiserVideo conditioning=false, Text conditioning=false2025.01 | 0.2 | 66.6 | 1.17 |