Audio Generation on MAESTRO (test)
0.131FAD (unconditional)Self-attention
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Self-attentionSub-decoder type=Self-attention2024.08 | 0.131 | 0.186 | 0.074 | 4.353 | |
| FlattenArchitecture=Flatten2024.08 | 0.14 | 0.176 | 0.068 | 4.482 | |
| Cross-attentionSub-decoder type=Cross-attention2024.08 | 0.145 | 0.19 | 0.065 | 4.314 | |
| NMTArchitecture=Nested Music Transformer2024.08 | 0.165 | 0.198 | 0.067 | 4.318 | |
| ParallelArchitecture=Parallel2024.08 | 0.166 | 0.206 | 0.075 | 4.669 | |
| DelayArchitecture=Delay2024.08 | 0.168 | 0.188 | 0.066 | 4.564 |