Visual Acoustic Matching on SoundSpaces-Speech unseen environments (test)
0.079RTE (s)LeMARA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LeMARATrain dataset=SoundSpaces-Speech, visual_features=true2023.07 | 0.079 | 0.69 | 1.53 | |
| AVITARTrain dataset=SoundSpaces-Speech2023.07 | 0.08 | 2.471 | 1.629 | |
| ViGASTrain dataset=SoundSpaces-Speech2023.07 | 0.108 | 4.373 | 1.232 | |
| LeMARA (no vis)Train dataset=SoundSpaces-Speech, visual_features=false2023.07 | 0.152 | 5.612 | 1.884 | |
| Input audioTrain dataset=SoundSpaces-Speech2023.07 | 0.32 | 1.427 | 1.274 |