Multi-talker Automatic Speech Recognition on Libri2Mix clean (test)
3.1WERID-30.D
Evaluation Results
| Method | Links | |
|---|---|---|
| ID-30.DLLM usage=With LLMs2026.03 | 3.1 | |
| SOP-Llama-3BLLM usage=With LLMs, Backbone=Llama-3B, Protocol=SOP2026.03 | 3.6 | |
| TUnetASR Backend=Whisper V3, N for BoN=4, Selection Strategy=Biometric2026.07 | 3.84 | |
| ID-20.DLLM usage=With LLMs2026.03 | 3.9 | |
| SOP-Llama-1BLLM usage=With LLMs, Backbone=Llama-1B, Protocol=SOP2026.03 | 4 | |
| SOT-Llama-3BLLM usage=With LLMs, Backbone=Llama-3B, Protocol=SOT2026.03 | 4.1 | |
| SOT-Llama-1BLLM usage=With LLMs, Backbone=Llama-1B, Protocol=SOT2026.03 | 4.6 | |
| ConvUnetASR Backend=Whisper V3, N for BoN=4, Selection Strategy=Biometric2026.07 | 4.99 | |
| SepReformerASR Backend=Whisper V3, Processing=full2026.07 | 5.02 | |
| GEncSep w/o decoderLLM usage=Without LLMs, SSL for speech encoder=true, Alignment protocol=serialized CTC2026.03 | 6.1 | |
| UMELLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 6.4 | |
| GEncSepLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 6.6 | |
| Training from ScratchLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 7 | |
| WavLM-CLNLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 7.6 | |
| C-HuBERT LARGELLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 7.8 | |
| W2V-Sidecar-ft.LLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 8.1 | |
| HCMLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 8.2 | |
| MeanFlow-TSEASR Backend=Whisper V3, Processing=full2026.07 | 9.05 | |
| MUSE-TSASRLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 11 | |
| SepReformerASR Backend=Whisper V3, Processing=chunk2026.07 | 12.9 | |
| SoloSpeechLLM usage=Without LLMs, SSL for speech encoder=true2026.03 | 15 |