Audio-visual generation on MultiDialog (test)
0.624SIMProposed
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ProposedSystem Type=Audio-Visual Spoken Dialogue System2024.06 | 0.624 | 30.323 | 7.298 | 7.39 | |
| AVSR + LM + TTS + TFGSystem Type=Cascaded System2024.06 | 0.433 | 30.581 | 7.041 | 7.64 | |
| d-GSLMSystem Type=Spoken Dialogue System2024.06 | 0.211 | — | — | — | |
| SpeechGPTSystem Type=Spoken Dialogue System2024.06 | 0.194 | — | — | — |