Text-to-Audio Generation on MusicCaps
108.69FDopenl3Stable Audio
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Stable Audiochannels/sr=2/44.1kHz, output length=95 sec, inference time=8 sec2024.02 | 108.69 | 0.8 | 0.46 | |
| Stable Audio w/ CLAPourschannels/sr=2/44.1kHz, output length=23 sec, inference time=4 sec2024.02 | 118.09 | 0.97 | 0.44 | |
| Stable Audio w/ CLAPLAIONchannels/sr=2/44.1kHz, output length=23 sec, inference time=4 sec2024.02 | 123.3 | 1.09 | 0.43 | |
| Stable Audio w/ T5channels/sr=2/44.1kHz, output length=23 sec, inference time=4 sec2024.02 | 126.93 | 1.06 | 0.41 | |
| MusicGen-largechannels/sr=1/32kHz, output length=95 sec, inference time=242 sec2024.02 | 197.12 | 0.85 | 0.36 | |
| MusicGen-smallchannels/sr=1/32kHz, output length=95 sec, inference time=126 sec2024.02 | 205.65 | 0.96 | 0.33 | |
| MusicGen-large-stereochannels/sr=2/32kHz, output length=95 sec, inference time=295 sec2024.02 | 216.07 | 1.04 | 0.32 | |
| AudioLDM2-48kHzchannels/sr=1/48kHz, output length=95 sec, inference time=242 sec2024.02 | 299.47 | 2.77 | 0.22 | |
| AudioLDM2-largechannels/sr=1/16kHz, output length=95 sec, inference time=37 sec2024.02 | 339.25 | 1.46 | 0.3 | |
| AudioLDM2-musicchannels/sr=1/16kHz, output length=95 sec, inference time=38 sec2024.02 | 354.05 | 1.53 | 0.3 |