Transcription on LibriTTS + DEMAND mixtures Background
2.2WERQwen2-Audio
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2-AudioSystem Setup=Upper Bound: Oracle Speaker + Auditory LLM2025.02 | 2.2 | 94.6 | |
| AAD-LLMSystem Setup=Intention-Informed AAD-LLM2025.02 | 10.8 | 86.1 | |
| Qwen2-AudioSystem Setup=Proposed Baseline: Extracted Speaker + Auditory LLM2025.02 | 14.3 | 82.1 | |
| Qwen2-AudioSystem Setup=Lower Bound: Random Speaker + Auditory LLM2025.02 | 57.7 | 45.2 | |
| Qwen2-AudioSystem Setup=Auditory LLM without Intention2025.02 | 80.4 | 27.9 | |
| Qwen-AudioSystem Setup=Auditory LLM without Intention2025.02 | 82.8 | 20.7 | |
| SALMONNSystem Setup=Auditory LLM without Intention2025.02 | 98.1 | 19.5 | |
| LTU-ASSystem Setup=Auditory LLM without Intention2025.02 | 118.5 | 27.4 |