Automatic Speech Recognition on LibriSpeech 100h clean (dev)
2.1WERSpeechT5
Evaluation Results
| Method | Links | |
|---|---|---|
| SpeechT5LM=Transformer2021.10 | 2.1 | |
| wav2vec 2.0 BASELM=Transformer2021.10 | 2.2 | |
| BaselineLM=Transformer2021.10 | 2.3 | |
| wav2vec 2.0 BASELM=4-gram2021.10 | 2.7 | |
| HuBERT BASELM=4-gram2021.10 | 2.7 | |
| NST with LM FusionSupervision=Semi-supervised, NST=true, LM Fusion=true2020.05 | 3.9 | |
| DiscreteBERTLM=4-gram2021.10 | 4 | |
| Whisper-M + Llama 8B (4bit, r=4)Train=Joint, Decode=PF-S2T2026.06 | 4.2 | |
| NST before LM FusionSupervision=Semi-supervised, NST=true, LM Fusion=false2020.05 | 4.3 | |
| SpeechT5Joint CTC/Attention=true2021.10 | 4.3 | |
| Whisper-M + Llama 8B (4bit, r=4)Train=Joint, Decode=S2T2026.06 | 4.8 | |
| BaselineJoint CTC/Attention=true2021.10 | 4.9 | |
| Lüscher et al.Supervision=Supervised2020.05 | 5 | |
| Whisper-M + Llama 8B (4bit, r=4)Train=PF-S2T, Decode=PF-S2T2026.06 | 5 | |
| Baseline (LAS + SpecAugment)Supervision=Supervised, SpecAugment=true2020.05 | 5.3 | |
| Hsu et al.Supervision=Semi-supervised (w/ LibriSpeech 860h)2020.05 | 5.39 | |
| SpeechT5Joint CTC/Attention=false2021.10 | 5.4 | |
| Kahn et al.Supervision=Semi-supervised (w/ LibriSpeech 860h)2020.05 | 5.41 | |
| HuBERT BASE2021.10 | 5.5 | |
| BaselineJoint CTC/Attention=false2021.10 | 5.8 | |
| Whisper-M + Llama 8B (4bit, r=4)Train=S2T, Decode=S2T2026.06 | 5.9 | |
| wav2vec 2.0 BASE2021.10 | 6.1 | |
| Whisper-MTrain=–, Decode=Beam2026.06 | 6.5 | |
| Kahn et al.Supervision=Supervised2020.05 | 7.78 | |
| Hsu et al.Supervision=Supervised2020.05 | 14 |