Streaming Automatic Speech Recognition on STOP1
12.09WERStreaming CTC-WS
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Streaming CTC-WSDecoder=RNN-T, Chunk Size=1120 ms2026.05 | 12.09 | 88.67 | 91.51 | 86 | |
| GPU-PBDecoder=RNN-T, Chunk Size=1120 ms2026.05 | 12.21 | 84.88 | 95.91 | 76.12 | |
| Streaming CTC-WSDecoder=CTC, Chunk Size=1120 ms2026.05 | 12.83 | 89.61 | 93.87 | 85.72 | |
| GPU-PBDecoder=CTC, Chunk Size=1120 ms2026.05 | 15.42 | 78.78 | 96.39 | 66.62 | |
| Non-biasing BaselineDecoder=RNN-T, Chunk Size=1120 ms2026.05 | 15.46 | 74.08 | 96.89 | 59.96 | |
| Non-biasing BaselineDecoder=CTC, Chunk Size=1120 ms2026.05 | 18.36 | 66.84 | 96.87 | 51.02 |