Automatic Speech Recognition on LibriSpeech clean segmented (test)
1.7WEROurs.small
Evaluation Results
| Method | Links | |
|---|---|---|
| Ours.smallDecoder=Joint CTC-Attention, Chunk Size (seconds)=3.842025.12 | 1.7 | |
| Ours.smallDecoder=Joint CTC-Attention, Chunk Size (seconds)=0.962025.12 | 1.8 | |
| Ours.baseDecoder=Joint CTC-Attention, Chunk Size (seconds)=3.842025.12 | 1.9 | |
| Ours.smallDecoder=Attention Rescoring, Chunk Size (seconds)=0.962025.12 | 2 | |
| Ours.smallDecoder=Attention Decoding, Chunk Size (seconds)=0.962025.12 | 2 | |
| Ours.baseDecoder=Joint CTC-Attention, Chunk Size (seconds)=0.962025.12 | 2.1 | |
| Ours.baseDecoder=Attention Decoding, Chunk Size (seconds)=0.962025.12 | 2.2 | |
| Ours.baseDecoder=Attention Rescoring, Chunk Size (seconds)=0.962025.12 | 2.3 | |
| Whisper small.enDecoder=Attention Decoding, Chunk Size (seconds)=302025.12 | 3.2 | |
| Whisper base.enDecoder=Attention Decoding, Chunk Size (seconds)=302025.12 | 4.1 |