Audio-visual speech separation on VoxCeleb2 (test)
14.6SI-SNRiDolphin
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Dolphin2025.09 | 14.6 | — | 15.1 | 3.17 | |
| AV-Mossformer2results_source=reproduced using official code2025.09 | 14 | — | 14.6 | 3.13 | |
| IIANet2023.08 | 13.6 | — | 14.3 | 3.12 | |
| IIANetresults_source=reported directly from original paper2025.09 | 13.6 | — | 14.3 | 3.12 | |
| Swift-Netresults_source=reproduced using official code2025.09 | 12.8 | — | 13.5 | 2.99 | |
| IIANet-fast2023.08 | 12.6 | — | 13.6 | 3.01 | |
| RTFS-Netresults_source=reported directly from original paper2025.09 | 12.4 | — | 13.6 | 3 | |
| CTCNet2023.08 | 11.9 | — | 13.1 | 3 | |
| CTCNetresults_source=reported directly from original paper2025.09 | 11.9 | — | 13.1 | 3 | |
| AVLIT-82023.08 | 9.4 | — | 9.9 | 2.23 | |
| AVLiT-8results_source=reported directly from original paper2025.09 | 9.4 | — | 9.9 | 2.23 | |
| Visualvoice2023.08 | 9.3 | — | 10.2 | 2.45 | |
| Visualvoiceresults_source=reported directly from original paper2025.09 | 9.3 | — | 10.2 | 2.45 | |
| AVConvTasNet2023.08 | 9.2 | — | 9.8 | 2.17 | |
| AV-ConvTasNetresults_source=reported directly from original paper2025.09 | 9.2 | — | 9.8 | 2.17 | |
| CaffNet-C2023.08 | 7.6 | — | — | — | |
| FaceFilterinput=static face2021.01 | — | 2.53 | — | — | |
| VISUALVOICEinput=static face2021.01 | — | 7.21 | — | — | |
| VISUALVOICEinput=full audio-visual2021.01 | — | 10.2 | — | — |