Multi-talker Automatic Speech Recognition on Libri2Mix Clean (dev)
3WERID-30.D
Evaluation Results
| Method | Links | |
|---|---|---|
| ID-30.DLLM usage=With LLMs2026.03 | 3 | |
| SOP-Llama-3BLLM Usage=With LLMs2026.03 | 3.5 | |
| SOP-Llama-3BLLM usage=With LLMs, Backbone=Llama-3B, Protocol=SOP2026.03 | 3.5 | |
| ID-20.DLLM usage=With LLMs2026.03 | 3.8 | |
| SOP-Llama-1BLLM Usage=With LLMs2026.03 | 3.9 | |
| SOP-Llama-1BLLM usage=With LLMs, Backbone=Llama-1B, Protocol=SOP2026.03 | 3.9 | |
| SOT-Llama-3BLLM Usage=With LLMs2026.03 | 4 | |
| ID-8LLM Usage=With LLMs2026.03 | 4 | |
| SOT-Llama-3BLLM usage=With LLMs, Backbone=Llama-3B, Protocol=SOT2026.03 | 4 | |
| SOT-Llama-1BLLM Usage=With LLMs2026.03 | 4.6 | |
| SOT-Llama-1BLLM usage=With LLMs, Backbone=Llama-1B, Protocol=SOT2026.03 | 4.6 | |
| w/o decoder: CTCLLM Usage=Without LLMs, Encoder=SSL2026.03 | 6 | |
| GEncSep w/o decoderLLM usage=Without LLMs, SSL for speech encoder=true, Alignment protocol=serialized CTC2026.03 | 6 | |
| GEncSepLLM Usage=Without LLMs, Encoder=SSL2026.03 | 6.4 | |
| GEncSepLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 6.4 | |
| C-HuBERT LARGELLM Usage=Without LLMs, Encoder=SSL2026.03 | 6.6 | |
| C-HuBERT LARGELLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 6.6 | |
| Training from ScratchLLM Usage=Without LLMs, Encoder=SSL2026.03 | 6.8 | |
| Training from ScratchLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 6.8 | |
| WavLM-CLNLLM Usage=Without LLMs, Encoder=SSL2026.03 | 7.1 | |
| WavLM-CLNLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 7.1 | |
| W2V-Sidecar-ft.LLM Usage=Without LLMs, Encoder=SSL2026.03 | 7.7 | |
| W2V-Sidecar-ft.LLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 7.7 |