Transcription on LibriTTS + DEMAND mixtures Foreground
1.9WERQwen2-Audio
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2-AudioSystem Setup=Upper Bound: Oracle Speaker + Auditory LLM2025.02 | 1.9 | 94.8 | |
| AAD-LLMSystem Setup=Intention-Informed AAD-LLM2025.02 | 10.6 | 86.3 | |
| Qwen2-AudioSystem Setup=Proposed Baseline: Extracted Speaker + Auditory LLM2025.02 | 14 | 82.3 | |
| Qwen2-AudioSystem Setup=Lower Bound: Random Speaker + Auditory LLM2025.02 | 56.8 | 45.2 | |
| Qwen2-AudioSystem Setup=Auditory LLM without Intention2025.02 | 81 | 25.7 | |
| WavLLMSystem Setup=Auditory LLM without Intention2025.02 | 84.4 | 21 | |
| Qwen-AudioSystem Setup=Auditory LLM without Intention2025.02 | 96 | 18.5 | |
| SALMONNSystem Setup=Auditory LLM without Intention2025.02 | 122.1 | 25.5 | |
| LTU-ASSystem Setup=Auditory LLM without Intention2025.02 | 148 | 26.3 |