Automatic Speech Recognition on ViVOS (test)
11.96CERViSpeechFormer
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ViSpeechFormerDecoding Level=phoneme, Decoder Parameters=2,007,8922026.02 | 11.96 | 30.49 | |
| Conv-TransformerDecoding Level=subword, Decoder Parameters=2,559,0892026.02 | 16.23 | 32.69 | |
| Speech TransformerDecoding Level=character, Decoder Parameters=1,249,9202026.02 | 18.54 | 34.83 | |
| TASADecoding Level=subword, Decoder Parameters=2,605,0412026.02 | 21.1 | 34.7 | |
| ConformerDecoding Level=subword, Decoder Parameters=4,302,0322026.02 | 22.87 | 37.61 | |
| ZipFormerDecoding Level=subword, Decoder Parameters=4,302,0322026.02 | 26.34 | 38.87 | |
| Multi-ConvFormerDecoding Level=subword, Decoder Parameters=4,302,0322026.02 | 30.98 | 44.8 |