Music Generation on Song Describer Dataset (test)
331.7FDopenl3AudioLDM2 Large
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AudioLDM2 LargeModel Type=Diffusion, Text Prompt Conditioning=Flan-T5 embeddings2026.07 | 331.7 | — | — | |
| TangoFluxModel Type=Diffusion, Text Prompt Conditioning=Flan-T5 embeddings2026.07 | 235.6 | — | — | |
| MusicGen-Stereo-L♠Model Type=Autoregressive, Text Prompt Conditioning=T5 embeddings2026.07 | 228.9 | — | — | |
| MusicGen-large-stereochannels/sr=2/32kHz, output length=472024.07 | 190.47 | 0.52 | 0.31 | |
| Stable Audio 1.0channels/sr=2/44.1kHz, output length=95 sec2024.07 | 142.5 | 0.4 | 0.38 | |
| Stable Audio OpenModel Type=Diffusion, Text Prompt Conditioning=T5/Flan-T5 embeddings2026.07 | 138.6 | — | — | |
| Stable Audio Openchannels/sr=2/44.1kHz, output length=472024.07 | 96.51 | 0.55 | 0.41 | |
| ETTAModel Type=Diffusion, Text Prompt Conditioning=T5 embeddings2026.07 | 95.7 | — | — | |
| UALMModel Type=LLM, Text Prompt Conditioning=BPE Tokens2026.07 | 83.7 | — | — | |
| Stable Audio 2.0channels/sr=2/44.1kHz, output length=285 sec2024.07 | 81.05 | 0.39 | 0.42 | |
| Audex 2BModel Type=LLM, Text Prompt Conditioning=BPE Tokens2026.07 | 78.4 | — | — | |
| UALM-GenModel Type=Autoregressive, Text Prompt Conditioning=BPE Tokens2026.07 | 74.4 | — | — | |
| Stable Audio 2.0channels/sr=2/44.1kHz, output length=190 sec2024.07 | 71.25 | 0.37 | 0.42 | |
| Audex 30B-A3BModel Type=LLM, Text Prompt Conditioning=BPE Tokens2026.07 | 62.7 | — | — | |
| Magenta RealTime♠Model Type=Autoregressive, Text Prompt Conditioning=MusicCoCa embeddings2026.07 | 40.5 | — | — |