Singing Voice Audio Reconstruction on Singing Voice Audio Dataset (test)
7.886FADshallow diffusion model
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| shallow diffusion modelconditioning=ground-truth f0, volume, and content2026.01 | 7.886 | 0.003 | 0.0921 | 0.029 | 0.7811 | -33.636 | 2.881 | 0.33 | 23.616 | |
| PSOLA2026.01 | 19.248 | 0.0083 | 0.1607 | 0.8442 | 0.6434 | -33.407 | 41.776 | 6.06 | 39.059 | |
| WORLD2026.01 | 31.397 | 0.0271 | 0.2287 | 1.4766 | 3.573 | -47.836 | 36.555 | 13.26 | 97.625 | |
| Diff-Pitcherevaluation_mode=zero-shot2026.01 | 39.321 | 0.0333 | 0.274 | 0.6366 | 2.7519 | -42.938 | 27.439 | 14.07 | 37.775 | |
| SiFiGANevaluation_mode=zero-shot2026.01 | 49.048 | 0.0792 | 0.4382 | 0.5526 | 3.0687 | -34.073 | 20.608 | 9.43 | 54.754 | |
| CLPCNet2026.01 | 77.702 | 0.0861 | 0.3644 | 0.9067 | 3.3673 | -48.504 | 16.946 | 24.12 | 65.323 |