Music Generation on Song Describer Dataset no-singing
78.7FDopenl3Stable Audio 2 (pre-trained)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Stable Audio 2 (pre-trained)channels/sr=2/44.1kHz, output length=2m+, inference time=8s2024.04 | 78.7 | 0.36 | 0.39 | |
| Stable Audio 2 (pre-trained)channels/sr=2/44.1kHz, output length=3m 10s, inference time=8s2024.04 | 89.33 | 0.34 | 0.39 | |
| MusicGen-large-stereochannels/sr=2/32kHz, output length=2m, inference time=6m 38s2024.04 | 204.03 | 0.49 | 0.28 | |
| MusicGen-large-stereochannels/sr=2/32kHz, output length=3m 10s, inference time=9m 32s2024.04 | 213.76 | 0.5 | 0.28 |