Multi-speaker Automatic Speech Recognition on AliMeeting (eval)
25.56CERTemporal Interleave
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Temporal Interleave#Params=0.7B, mask=true2026.06 | 25.56 | 27.96 | 2.4 | |
| Time-wise Concatenation#Params=0.7B, mask=true2026.06 | 25.64 | 28.51 | 2.87 | |
| Feature-wise Concatenation#Params=0.7B, mask=true2026.06 | 26.19 | 29.68 | 3.49 | |
| Sensevoice-small#Params=230M2026.06 | 26.62 | — | — | |
| Semantic Feature Only#Params=0.7B, mask=false2026.06 | 26.66 | 31.11 | 4.45 | |
| Temporal Interleave#Params=0.7B, mask=false2026.06 | 28.07 | 30.77 | 2.7 | |
| Feature-wise Concatenation#Params=0.7B, mask=false2026.06 | 28.3 | 32.23 | 3.93 | |
| VibeVoice-ASR#Params=7B2026.06 | 31.38 | 39.2 | 7.82 | |
| Paraformer+3D speaker#Params=70M2026.06 | 31.8 | 36.39 | 4.59 | |
| Paraformer+DiariZen-large#Params=140M2026.06 | 31.8 | 36.09 | 4.29 |