Automatic Speech Recognition on CommonVoice segmented (test)
11.4WEROurs.small
Evaluation Results
| Method | Links | |
|---|---|---|
| Ours.smallDecoder=Joint CTC-Attention, Chunk Size (seconds)=3.842025.12 | 11.4 | |
| Ours.smallDecoder=Attention Decoding, Chunk Size (seconds)=0.962025.12 | 12.1 | |
| Ours.smallDecoder=Joint CTC-Attention, Chunk Size (seconds)=0.962025.12 | 12.4 | |
| Whisper small.enDecoder=Attention Decoding, Chunk Size (seconds)=302025.12 | 12.6 | |
| Ours.smallDecoder=Attention Rescoring, Chunk Size (seconds)=0.962025.12 | 13.4 | |
| Ours.baseDecoder=Joint CTC-Attention, Chunk Size (seconds)=3.842025.12 | 13.7 | |
| Ours.baseDecoder=Attention Decoding, Chunk Size (seconds)=0.962025.12 | 14.1 | |
| Ours.baseDecoder=Joint CTC-Attention, Chunk Size (seconds)=0.962025.12 | 14.4 | |
| Ours.baseDecoder=Attention Rescoring, Chunk Size (seconds)=0.962025.12 | 16.3 | |
| Whisper base.enDecoder=Attention Decoding, Chunk Size (seconds)=302025.12 | 17.5 |