Automatic Speech Recognition on YouTube
14.05WERChunkFormer
Evaluation Results
| Method | Links | |
|---|---|---|
| ChunkFormercontext size [latt, c, r]=[128, 64, 128]2025.02 | 14.05 | |
| ChunkFormercontext size [latt, c, r]=[128, 256, 128]2025.02 | 14.2 | |
| ChunkFormercontext size [latt, c, r]=[256, 128, 128]2025.02 | 14.22 | |
| ChunkFormerdecoding_mode=full-context2025.02 | 15 | |
| Efficient Conformerdecoding_mode=full-context2025.02 | 16.19 | |
| Squeezeformerdecoding_mode=full-context2025.02 | 16.83 | |
| Conformerdecoding_mode=full-context2025.02 | 16.89 |