Speech-to-Singing conversion on English (test)
2.512LSDGT (vocoder)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GT (vocoder)source=Ground Truth acoustic codes, vocoder=BigVGAN2024.06 | 2.512 | 0.988 | 4.52 | 4.47 | 4.13 | |
| SVPTsemantic_features=XLSR-53 12th layer2024.06 | 5.213 | 0.967 | 3.46 | 3.68 | 3.61 | |
| SVPTsemantic_features=wav2vec 2.0 18th layer2024.06 | 5.462 | 0.956 | 3.44 | 3.48 | 3.39 | |
| AlignSTSarchitecture=diffusion-based, features=rhythm and pitch modal fusion2024.06 | 5.519 | 0.941 | 3.45 | 3.47 | 3.41 | |
| Wu and Yang, 2020architecture=GAN-based2024.06 | 6.913 | 0.896 | 3.12 | 3.21 | 3.15 | |
| Parekh et al., 2020architecture=CNN-based, training_database=NHSS2024.06 | 8.045 | 0.842 | 3.02 | 2.99 | 3.01 |