Predicting Sounds from Video on Greatest Hits
0.21Loudness ErrorFull system of [37]
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Full system of [37]Architecture=CNN + LSTM2016.12 | 0.21 | 0.44 | 3.85 | 0.47 | |
| Supervised tessellationBackbone=VGG-19, Visual Representation=RNN-FV pooled2016.12 | 0.24 | 0.35 | 4.44 | 0.46 | |
| Unsupervised tessellationBackbone=VGG-19, Visual Representation=RNN-FV pooled2016.12 | 0.26 | 0.33 | 4.76 | 0.48 | |
| Local tessellationBackbone=VGG-19, Visual Representation=RNN-FV pooled2016.12 | 0.27 | 0.32 | 4.83 | 0.47 | |
| Appearance matching2016.12 | 0.35 | 0.18 | 6.09 | 0.36 |