Visual Acoustic Matching on AVSpeech-Rooms unseen environments (test)
0.071RTE (s)LeMARA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LeMARATrain dataset=AVSpeech-Rooms, visual_features=true2023.07 | 0.071 | 6.298 | 0.571 | |
| ViGASTrain dataset=AVSpeech-Rooms2023.07 | 0.109 | 7.007 | 0.662 | |
| AVITARTrain dataset=AVSpeech-Rooms2023.07 | 0.136 | 2.894 | 1.29 | |
| LeMARA (no vis)Train dataset=AVSpeech-Rooms, visual_features=false2023.07 | 0.137 | 6.256 | 0.705 | |
| Input audioTrain dataset=AVSpeech-Rooms2023.07 | 0.31 | 1.327 | 2.107 |