Streaming Automatic Speech Recognition on STOP2
9.66WERStreaming CTC-WS
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Streaming CTC-WSDecoder=RNN-T, Chunk Size=1120 ms2026.05 | 9.66 | 95.6 | 97.47 | 93.79 | |
| GPU-PBDecoder=RNN-T, Chunk Size=1120 ms2026.05 | 10.18 | 93.62 | 96.83 | 90.62 | |
| Non-biasing BaselineDecoder=RNN-T, Chunk Size=1120 ms2026.05 | 10.43 | 92.16 | 98.6 | 86.51 | |
| Streaming CTC-WSDecoder=CTC, Chunk Size=1120 ms2026.05 | 10.48 | 95.06 | 97.35 | 92.88 | |
| GPU-PBDecoder=CTC, Chunk Size=1120 ms2026.05 | 11.86 | 90.53 | 98.05 | 84.09 | |
| Non-biasing BaselineDecoder=CTC, Chunk Size=1120 ms2026.05 | 12.09 | 88.26 | 98.76 | 79.79 |