Multi-speaker Automatic Speech Recognition on AliMeeting (test)
16.05cpCERSpeakerLM (7639h)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SpeakerLM (7639h)Parms=7B, Training Data / Hours=7639h2026.04 | 16.05 | — | — | — | — | — | |
| SpeakerLM* (7638 hours)#Params=7B2026.06 | 16.05 | — | — | — | 13.97 | 2.08 | |
| DM-ASR (S2SND) (CN 1300h)Parms=1.7B, Training Data / Hours=CN 1300h, Diarization Source=S2SND, Time Prediction Source=LLM-predicted, Label Perturbation Status=None2026.04 | 19.15 | 10.09 | — | 19.45 | — | — | |
| DM-ASR (DiariZen) (CN 1300h)Parms=1.7B, Training Data / Hours=CN 1300h, Diarization Source=DiariZen, Time Prediction Source=LLM-predicted, Label Perturbation Status=None2026.04 | 20.95 | 10.92 | — | 21.55 | — | — | |
| DM-ASR (S2SND) (CN+EN 2900h)Parms=1.7B, Training Data / Hours=CN+EN 2900h, Diarization Source=S2SND, Time Prediction Source=LLM-predicted, Label Perturbation Status=None2026.04 | 21.4 | 10.09 | 2.23 | 21.79 | — | — | |
| DM-ASR (DiariZen) (CN 1300h)Parms=0.6B, Training Data / Hours=CN 1300h, Diarization Source=DiariZen, Time Prediction Source=LLM-predicted, Label Perturbation Status=None2026.04 | 21.6 | 10.97 | — | 22.23 | — | — | |
| DM-ASR (DiariZen) (CN+EN 2900h)Parms=0.6B, Training Data / Hours=CN+EN 2900h, Diarization Source=DiariZen, Time Prediction Source=LLM-predicted, Label Perturbation Status=None2026.04 | 21.73 | 10.94 | — | 22.35 | — | — | |
| Paraformer+3D speaker*#Params=70M2026.06 | 23.2 | — | — | — | 21.3 | 1.9 | |
| DM-ASR (DiariZen) (CN+EN 2900h)Parms=1.7B, Training Data / Hours=CN+EN 2900h, Diarization Source=DiariZen, Time Prediction Source=LLM-predicted, Label Perturbation Status=None2026.04 | 23.32 | 10.92 | — | 24.02 | — | — | |
| DM-ASR (DiariZen) (CN 630h)Parms=0.6B, Training Data / Hours=CN 630h, Diarization Source=DiariZen, Time Prediction Source=LLM-predicted, Label Perturbation Status=None2026.04 | 23.46 | 11 | — | 24.09 | — | — | |
| Temporal Interleave#Params=0.7B, mask=true2026.06 | 27.16 | — | — | — | 23.61 | 3.55 | |
| SpeakerLM (2270h)Parms=7B, Training Data / Hours=2270h2026.04 | 27.97 | — | — | — | — | — | |
| VibeVoice-ASR(>9400h)Parms=7B, Training Data / Hours=>9400h2026.04 | 29.33 | 10.92 | — | 29.51 | — | — | |
| SpeakerLM (694h)Parms=7B, Training Data / Hours=694h2026.04 | 29.6 | — | — | — | — | — | |
| Feature-wise Concatenation#Params=0.7B, mask=true2026.06 | 29.64 | — | — | — | 24.27 | 5.37 | |
| Temporal Interleave#Params=0.7B, mask=false2026.06 | 29.71 | — | — | — | 26.22 | 3.49 | |
| Time-wise Concatenation#Params=0.7B, mask=true2026.06 | 30.1 | — | — | — | 25.29 | 4.81 | |
| Feature-wise Concatenation#Params=0.7B, mask=false2026.06 | 31.47 | — | — | — | 25.6 | 5.87 | |
| Semantic Feature Only#Params=0.7B, mask=false2026.06 | 31.94 | — | — | — | 26.12 | 5.82 | |
| SpeakerLM* (212 hours)#Params=7B2026.06 | 32.22 | — | — | — | 18.63 | 13.59 | |
| Paraformer+3D speaker#Params=70M2026.06 | 32.46 | — | — | — | 27.78 | 4.68 | |
| Gemini-3.0-ProParms=-2026.04 | 32.84 | 38.75 | — | 65.61 | — | — | |
| Paraformer+DiariZen-large#Params=140M2026.06 | 33.09 | — | — | — | 27.78 | 5.31 | |
| Tagspeech(103h)Parms=7B, Training Data / Hours=103h2026.04 | 33.84 | 22.13 | — | — | — | — | |
| Semi-end-to-end MS-ASR(1000+ h)Parms=3B, Training Data / Hours=1000+ h2026.04 | 35.1 | — | — | 36.36 | — | — | |
| VibeVoice-ASR#Params=7B2026.06 | 35.86 | — | — | — | 29.47 | 6.39 | |
| DiariZen+Whisper-large-v3Parms=1.5B2026.04 | 41.05 | 10.8 | — | 43.75 | — | — | |
| Qwen2.5-Omni-7BParms=7B2026.04 | 41.23 | 37.42 | — | — | — | — | |
| Gemini-2.5-ProParms=-2026.04 | 41.64 | 31.6 | — | 53.49 | — | — | |
| Pyannote+Whisper-large-v3Parms=1.5B2026.04 | 46.56 | 26.13 | — | — | — | — | |
| Sensevoice-small#Params=230M2026.06 | — | — | — | — | 25.08 | — |