Speech Recognition on LibriSpeech (test)
0.0133WERAdam
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| AdamArchitecture=Conformer (600M), Pre-training=w2v-BERT, Fine-tuning protocol=Semi-supervised (LibriLight), Optimizer=Adam2024.05 | 0.0133 | — | — | — | — | — | |
| FAdamArchitecture=Conformer (600M), Pre-training=w2v-BERT, Fine-tuning protocol=Semi-supervised (LibriLight), Optimizer=FAdam2024.05 | 0.0134 | — | — | — | — | — | |
| Adam (w2v-BERT paper)Architecture=Conformer (600M), Pre-training=w2v-BERT, Optimizer=Adam, Source=[53]2024.05 | 0.014 | — | — | — | — | — | |
| Gen3 Conformer XXLUnlabeled data (hrs)=60k, LM Fusion=Yes2020.10 | 0.014 | — | — | — | — | — | |
| Gen3 Conformer XXL+Unlabeled data (hrs)=60k, LM Fusion=Yes2020.10 | 0.014 | — | — | — | — | — | |
| Pre-trained Conformer XLUnlabeled data (hrs)=60k, LM Fusion=Yes2020.10 | 0.015 | — | — | — | — | — | |
| Pre-trained Conformer XXLUnlabeled data (hrs)=60k, LM Fusion=Yes2020.10 | 0.015 | — | — | — | — | — | |
| Gen3 Conformer XXLUnlabeled data (hrs)=60k, LM Fusion=No2020.10 | 0.015 | — | — | — | — | — | |
| Gen3 Conformer XXL+Unlabeled data (hrs)=60k, LM Fusion=No2020.10 | 0.015 | — | — | — | — | — | |
| Pre-trained Conformer XXLUnlabeled data (hrs)=60k, LM Fusion=No2020.10 | 0.016 | — | — | — | — | — | |
| Gen4 ContextNetUnlabeled data (hrs)=60k, LM Fusion=No2020.10 | 0.017 | — | — | — | — | — | |
| Gen4 ContextNetUnlabeled data (hrs)=60k, LM Fusion=Yes2020.10 | 0.017 | — | — | — | — | — | |
| Gen4 Conformer LUnlabeled data (hrs)=60k, LM Fusion=No2020.10 | 0.017 | — | — | — | — | — | |
| Gen4 Conformer LUnlabeled data (hrs)=60k, LM Fusion=Yes2020.10 | 0.017 | — | — | — | — | — | |
| Pre-trained Conformer XLUnlabeled data (hrs)=60k, LM Fusion=No2020.10 | 0.017 | — | — | — | — | — | |
| Pre-trained CTCUnlabeled data (hrs)=60k, LM Fusion=Yes2020.10 | 0.018 | — | — | — | — | — | |
| Conformer LUnlabeled data (hrs)=None, LM Fusion=No2020.10 | 0.021 | — | — | — | — | — | |
| Pre-trained CTCUnlabeled data (hrs)=60k, LM Fusion=No2020.10 | 0.022 | — | — | — | — | — | |
| ClozeGERBackbone=LLaMA-2-7b, Source Speech=false, LoRA=true, Logits Calibration=true, Post-processing=true2024.05 | 0.024 | — | — | 11.1 | — | — | |
| ClozeGERBackbone=SpeechGPT, Source Speech=true, LoRA=true, Logits Calibration=true, Post-processing=true2024.05 | 0.025 | — | — | 7.4 | — | — | |
| Whisper Baseline2024.05 | 0.027 | — | — | — | — | — | |
| CleanBackbone=Whisper-medium, SNR=30 dB2026.01 | 0.0299 | — | — | — | 21.84 | — | |
| CleanBackbone=Whisper-large, SNR=30 dB2026.01 | 0.0336 | — | — | — | 21.84 | — | |
| ASR + LMSystem=ASR + LM2025.12 | 0.0356 | — | — | — | — | — | |
| ASR + DLMSystem=ASR + DLM2025.12 | 0.0358 | — | — | — | — | — | |
| CleanBackbone=Whisper-small, SNR=30 dB2026.01 | 0.0364 | — | — | — | 21.84 | — | |
| CleanBackbone=Whisper-base, SNR=30 dB2026.01 | 0.0377 | — | — | — | 21.77 | — | |
| SOT + Local MoLEParm.(M)=35.092026.07 | 0.038 | — | — | — | — | — | |
| SOT-SACTCParm.(M)=35.772026.07 | 0.038 | — | — | — | — | — | |
| H-SAGEParm.(M)=35.752026.07 | 0.038 | — | — | — | — | — | |
| GLAD-SOTParm.(M)=35.182026.07 | 0.039 | — | — | — | — | — | |
| ASR onlySystem=ASR only2025.12 | 0.0426 | — | — | — | — | — | |
| SOTParm.(M)=36.072026.07 | 0.045 | — | — | — | — | — | |
| CleanBackbone=Whisper-tiny, SNR=30 dB2026.01 | 0.0666 | — | — | — | 21.84 | — | |
| Original2025.03 | 0.1108 | — | — | — | — | — | |
| Finetune2025.03 | 0.1367 | — | — | — | — | — | |
| OrthoGrad2025.03 | 0.1398 | — | — | — | — | — | |
| Full precisionBits=32, Model Architecture=Conformer2026.03 | 0.1594 | — | — | — | — | 4.84 | |
| ESCBits=8, Model Architecture=Conformer2026.03 | 0.1601 | — | — | — | — | 4.83 | |
| PercentileBits=8, Model Architecture=Conformer2026.03 | 0.1607 | — | — | — | — | 4.85 | |
| MSEBits=8, Model Architecture=Conformer2026.03 | 0.1609 | — | — | — | — | 4.87 | |
| EntropyBits=8, Model Architecture=Conformer2026.03 | 0.1627 | — | — | — | — | 4.93 | |
| MaxBits=8, Model Architecture=Conformer2026.03 | 0.1654 | — | — | — | — | 5.04 | |
| GDR-GMA2025.03 | 0.3252 | — | — | — | — | — | |
| ESCBits=4, Model Architecture=Conformer2026.03 | 0.3849 | — | — | — | — | 13.5 | |
| MSEBits=4, Model Architecture=Conformer2026.03 | 0.4122 | — | — | — | — | 14.61 | |
| SlothSpeechBackbone=Whisper-medium, SNR=30 dB2026.01 | 0.4175 | — | — | — | 77.43 | — | |
| VMI-FGSMBackbone=Whisper-large, SNR=30 dB2026.01 | 0.4275 | — | — | — | 22.07 | — | |
| MI-FGSMBackbone=Whisper-large, SNR=30 dB2026.01 | 0.4535 | — | — | — | 22.11 | — | |
| SAGOBackbone=Whisper-large, SNR=30 dB2026.01 | 0.4567 | — | — | — | 21.75 | — | |
| PGDBackbone=Whisper-large, SNR=30 dB2026.01 | 0.4793 | — | — | — | 21.92 | — | |
| SlothSpeechBackbone=Whisper-large, SNR=30 dB2026.01 | 0.4834 | — | — | — | 78.65 | — | |
| PercentileBits=4, Model Architecture=Conformer2026.03 | 0.5083 | — | — | — | — | 18.6 | |
| EntropyBits=4, Model Architecture=Conformer2026.03 | 0.5083 | — | — | — | — | 18.76 | |
| SlothSpeechBackbone=Whisper-small, SNR=30 dB2026.01 | 0.5645 | — | — | — | 102.15 | — | |
| SlothSpeechBackbone=Whisper-tiny, SNR=30 dB2026.01 | 0.6006 | — | — | — | 123.93 | — | |
| MOREBackbone=Whisper-large, SNR=30 dB2026.01 | 0.609 | — | — | — | 277.65 | — | |
| SlothSpeechBackbone=Whisper-base, SNR=30 dB2026.01 | 0.6933 | — | — | — | 152.82 | — | |
| MI-FGSMBackbone=Whisper-medium, SNR=30 dB2026.01 | 0.7252 | — | — | — | 27.54 | — | |
| VMI-FGSMBackbone=Whisper-medium, SNR=30 dB2026.01 | 0.7366 | — | — | — | 27.02 | — | |
| SAGOBackbone=Whisper-medium, SNR=30 dB2026.01 | 0.7467 | — | — | — | 27 | — | |
| MOREBackbone=Whisper-medium, SNR=30 dB2026.01 | 0.7784 | — | — | — | 202.34 | — | |
| PGDBackbone=Whisper-medium, SNR=30 dB2026.01 | 0.7877 | — | — | — | 28.87 | — | |
| VMI-FGSMBackbone=Whisper-small, SNR=30 dB2026.01 | 0.8315 | — | — | — | 28.06 | — | |
| MI-FGSMBackbone=Whisper-small, SNR=30 dB2026.01 | 0.8349 | — | — | — | 28.37 | — | |
| SAGOBackbone=Whisper-small, SNR=30 dB2026.01 | 0.8394 | — | — | — | 28.28 | — | |
| PGDBackbone=Whisper-small, SNR=30 dB2026.01 | 0.8583 | — | — | — | 31.4 | — | |
| NegGrad+2025.03 | 0.859 | — | — | — | — | — | |
| MOREBackbone=Whisper-small, SNR=30 dB2026.01 | 0.8615 | — | — | — | 238.52 | — | |
| MI-FGSMBackbone=Whisper-base, SNR=30 dB2026.01 | 0.9308 | — | — | — | 33.22 | — | |
| VMI-FGSMBackbone=Whisper-tiny, SNR=30 dB2026.01 | 0.9323 | — | — | — | 37.53 | — | |
| MOREBackbone=Whisper-base, SNR=30 dB2026.01 | 0.937 | — | — | — | 324.05 | — | |
| SAGOBackbone=Whisper-base, SNR=30 dB2026.01 | 0.9372 | — | — | — | 32.95 | — | |
| PGDBackbone=Whisper-base, SNR=30 dB2026.01 | 0.939 | — | — | — | 33.21 | — | |
| MOREBackbone=Whisper-tiny, SNR=30 dB2026.01 | 0.9473 | — | — | — | 300.79 | — | |
| SAGOBackbone=Whisper-tiny, SNR=30 dB2026.01 | 0.9564 | — | — | — | 34.44 | — | |
| VMI-FGSMBackbone=Whisper-base, SNR=30 dB2026.01 | 0.9608 | — | — | — | 37.72 | — | |
| MI-FGSMBackbone=Whisper-tiny, SNR=30 dB2026.01 | 0.9618 | — | — | — | 34.37 | — | |
| PGDBackbone=Whisper-tiny, SNR=30 dB2026.01 | 0.9622 | — | — | — | 35.3 | — | |
| SCRUB2025.03 | 1 | — | — | — | — | — | |
| MaxBits=4, Model Architecture=Conformer2026.03 | 1.4414 | — | — | — | — | 84.81 | |
| B2T connectionLayer configuration=Enc-Dec: 6L-6L2022.06 | — | 0.0386 | 0.0894 | — | — | — | |
| B2T connectionLayer configuration=Enc-Dec: 12L-6L2022.06 | — | 0.0348 | 0.0768 | — | — | — | |
| Centaurus Base (full SSM)Protocol=Online (streaming), no attention, Parameters (M)=12.4, FLOPs (G)=20.62025.01 | — | 0.06 | 0.131 | — | — | — | |
| Centaurus with causal conv-blockProtocol=Online (streaming), no attention, Parameters (M)=23.6, FLOPs (G)=43.72025.01 | — | 0.048 | 0.106 | — | — | — | |
| Centaurus with FFNProtocol=Online (streaming), no attention, Parameters (M)=18, FLOPs (G)=32.12025.01 | — | 0.054 | 0.115 | — | — | — | |
| Centaurus with Mamba macro-blockProtocol=Online (streaming), no attention, Parameters (M)=29.9, FLOPs (G)=46.92025.01 | — | 0.044 | 0.102 | — | — | — | |
| ConformerProtocol=Offline (full-context), Parameters (M)=30.7, FLOPs (G)=45.22025.01 | — | 0.02 | 0.043 | — | — | — | |
| ConformerProtocol=Online (streaming), Parameters (M)=30.72025.01 | — | 0.046 | 0.099 | — | — | — | |
| ContextNetProtocol=Offline (full-context), Parameters (M)=31.42025.01 | — | 0.024 | 0.054 | — | — | — | |
| ContextNetProtocol=Online (streaming), Parameters (M)=31.42025.01 | — | 0.045 | 0.1 | — | — | — | |
| Full ConvProtocol=Offline (full-context)2025.01 | — | 0.033 | 0.105 | — | — | — | |
| Post-LNLayer configuration=Enc-Dec: 6L-6L2022.06 | — | 0.0419 | 0.0874 | — | — | — | |
| Pre-LNLayer configuration=Enc-Dec: 6L-6L2022.06 | — | 0.0422 | 0.0965 | — | — | — | |
| Pre-LNLayer configuration=Enc-Dec: 12L-6L2022.06 | — | 0.0349 | 0.0822 | — | — | — | |
| TransformerProtocol=Offline (full-context), Parameters (M)=292025.01 | — | 0.031 | 0.073 | — | — | — | |
| TransformerProtocol=Online (streaming), Parameters (M)=18.92025.01 | — | 0.05 | 0.116 | — | — | — | |
| Wav2Vec2Protocol=Offline (full-context), Parameters (M)=94.4, FLOPs (G)=2852025.01 | — | 0.021 | 0.048 | — | — | — |