Multi-speaker Automatic Speech Recognition on Aishell4 (eval)
17.17Character Error Rate (CER)SpeakerLM* (7638 hours)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SpeakerLM* (7638 hours)#Params=7B2026.06 | 17.17 | 18.37 | 1.2 | |
| Temporal Interleave#Params=0.7B, mask=true2026.06 | 17.18 | 19.98 | 2.8 | |
| Feature-wise Concatenation#Params=0.7B, mask=true2026.06 | 17.54 | 22.24 | 4.7 | |
| Time-wise Concatenation#Params=0.7B, mask=true2026.06 | 17.56 | 20.95 | 3.39 | |
| SpeakerLM* (212 hours)#Params=7B2026.06 | 17.75 | 26.14 | 8.39 | |
| Semantic Feature Only#Params=0.7B, mask=false2026.06 | 18.38 | 23.08 | 4.7 | |
| Temporal Interleave#Params=0.7B, mask=false2026.06 | 18.41 | 21.45 | 3.04 | |
| Feature-wise Concatenation#Params=0.7B, mask=false2026.06 | 18.73 | 23.54 | 4.81 | |
| Sensevoice-small#Params=230M2026.06 | 18.86 | — | — | |
| VibeVoice-ASR#Params=7B2026.06 | 21.65 | 26.23 | 4.58 | |
| Paraformer+3D speaker#Params=70M2026.06 | 22.67 | 28.29 | 5.62 | |
| Paraformer+DiariZen-large#Params=140M2026.06 | 22.67 | 26.34 | 3.67 | |
| Paraformer+3D speaker*#Params=70M2026.06 | 23.02 | 26.01 | 2.99 |