Automatic Speech Recognition on LibriSpeech clean
1.16WERERNIE 5.0
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ERNIE 5.02026.02 | 1.16 | — | — | — | |
| Samba-ASRorganization=SandLogic2025.01 | 1.17 | — | — | — | |
| Qwen3-Omni-Instruct2026.02 | 1.22 | — | — | — | |
| Kimi Audio2026.02 | 1.28 | — | — | — | |
| StepAudio 2.5 ASRMTP training=true2026.05 | 1.38 | — | — | — | |
| Kimi Audio2026.04 | 1.38 | — | — | — | |
| GPT-4o-Audio2026.02 | 1.39 | — | — | — | |
| GPT-4o-Transcribe2026.01 | 1.39 | — | — | — | |
| nvidia/parakeet-tdt-1.1bparameters=1.1b2025.01 | 1.4 | — | — | — | |
| StepAudio 2.5 ASRMTP training=false2026.05 | 1.4 | — | — | — | |
| Step-Audio Flamingo 32026.04 | 1.4 | — | — | — | |
| Raon-Speech2026.04 | 1.44 | — | — | — | |
| ERNIE 5.0-Base2026.02 | 1.47 | — | — | — | |
| nvidia/canary-1bparameters=1b2025.01 | 1.48 | — | — | — | |
| Whisper-large-v32026.01 | 1.51 | — | — | — | |
| MiniCPM-o 4.52026.04 | 1.51 | — | — | — | |
| LongCat-Flash-Omni2026.02 | 1.57 | — | — | — | |
| Audio Flamingo 3Backbone=Audio Flamingo 3, Method=Baseline2026.06 | 1.57 | — | — | — | |
| Qwen2-Audio2026.03 | 1.6 | — | 1.6 | — | |
| Fun-Audio Chat2026.04 | 1.6 | — | — | — | |
| Mixed Training + Length Aug + Timestamp RegLength Augmentation=true, Timestamp Regularization=true2026.04 | 1.62 | — | — | — | |
| Qwen3-ASR-1.7BParameters=1.7B2026.01 | 1.63 | — | — | — | |
| DAPlin=✓, ldeep=✓, FRR=14.91, Pre. GFLOPs=566.30, FLOPs Ratio=72.52, Setting=Aggressive (τin=0.80, τdeep=0.70)2026.04 | 1.63 | — | — | — | |
| APinlin=✓, ldeep=-, FRR=78.64, Pre. GFLOPs=612.93, FLOPs Ratio=78.49, Setting=Aggressive (τin=0.80, τdeep=0.70)2026.04 | 1.64 | — | — | — | |
| DAPlin=✓, ldeep=✓, FRR=33.76, Pre. GFLOPs=686.39, FLOPs Ratio=87.89, Setting=Conservative (τin=0.90, τdeep=0.80)2026.04 | 1.64 | — | — | — | |
| Mixed Training + Length Aug + Reduced TFLength Augmentation=true, Reduced Teacher Forcing=true2026.04 | 1.64 | — | — | — | |
| Vanillalin=-, ldeep=-, FRR=100.0, Pre. GFLOPs=780.94, FLOPs Ratio=100.0, Setting=N/A2026.04 | 1.65 | — | — | — | |
| APdeeplin=-, ldeep=✓, FRR=14.30, Pre. GFLOPs=718.12, FLOPs Ratio=91.96, Setting=Aggressive (τin=0.80, τdeep=0.70)2026.04 | 1.65 | — | — | — | |
| APinlin=✓, ldeep=-, FRR=93.56, Pre. GFLOPs=730.28, FLOPs Ratio=93.51, Setting=Conservative (τin=0.90, τdeep=0.80)2026.04 | 1.65 | — | — | — | |
| APdeeplin=-, ldeep=✓, FRR=33.29, Pre. GFLOPs=731.96, FLOPs Ratio=93.73, Setting=Conservative (τin=0.90, τdeep=0.80)2026.04 | 1.66 | — | — | — | |
| FunASR-MLT-Nano2026.01 | 1.68 | — | — | — | |
| Qwen3-ASR-1.7B2026.05 | 1.69 | — | — | — | |
| CrisperWhisper2026.04 | 1.71 | — | — | — | |
| ASR-only Baseline2026.04 | 1.72 | — | — | — | |
| Mixed Training + Length AugLength Augmentation=true2026.04 | 1.72 | — | — | — | |
| Qwen2.5-Omni2026.04 | 1.73 | — | — | — | |
| Qwen2.5-Omni-7B + CoATBackbone=Qwen2.5-Omni-7B, Method=CoAT2026.06 | 1.77 | — | — | — | |
| Qwen2-Audio2026.01 | 1.8 | — | — | — | |
| MinMo2026.01 | 1.8 | — | — | — | |
| FunASR-Nano2026.05 | 1.8 | — | — | — | |
| Qwen2.5-Omni-7BBackbone=Qwen2.5-Omni-7B, Method=Baseline2026.06 | 1.8 | — | — | — | |
| Whisper-large-v3Model size=large2026.02 | 1.81 | — | — | — | |
| Mixed Training2026.04 | 1.81 | — | — | — | |
| nyrahealth/CrisperWhisper2025.01 | 1.82 | — | — | — | |
| Audio Flamingo 3 + CoATBackbone=Audio Flamingo 3, Method=CoAT2026.06 | 1.99 | — | — | — | |
| MoST2026.01 | 2 | — | — | — | |
| openai/whisper-large-v3model_size=large-v32025.01 | 2.01 | — | — | — | |
| Phi-4 Multimodal2026.01 | 2.1 | — | — | — | |
| Parakeet-TDT-v2Correction method=None2024.05 | 2.1 | — | — | — | |
| ECLMCorrection model size=0.5B, Decoding strategy=greedy, Base model=Parakeet-TDT-v22024.05 | 2.1 | — | — | 0 | |
| Qwen3-ASR-0.6BParameters=0.6B2026.01 | 2.11 | — | — | — | |
| OSUM2026.03 | 2.19 | — | 2.19 | — | |
| Qwen-Audio2026.04 | 2.19 | — | — | — | |
| Qwen3-ASR-1.7BModel Scale=1.7B2026.05 | 2.2 | — | — | — | |
| Fun-Audio Omni2026.04 | 2.28 | — | — | — | |
| HyperCLOVA X 8B Omni2026.04 | 2.28 | — | — | — | |
| VibeVoice-ASR2026.05 | 2.3 | — | — | — | |
| Qwen2-Audio + CoATBackbone=Qwen2-Audio, Method=CoAT2026.06 | 2.3 | — | — | — | |
| Step-AudioTrain audio duration(h)=>1000k2026.04 | 2.36 | — | — | — | |
| Qwen2.5-OmniTrain audio duration(h)=>1000k2026.04 | 2.37 | — | — | — | |
| AR Beam2026.06 | 2.4 | — | — | — | |
| Ark-Base+TD+OPDModel Scale=0.6B, TD=Teacher-data adaptation, OPD=On-policy distillation2026.05 | 2.45 | — | — | — | |
| UniAudio 2.02026.02 | 2.71 | — | — | — | |
| Gemini-3-Pro2026.02 | 2.74 | — | — | — | |
| Doubao-ASR2026.01 | 2.78 | — | — | — | |
| Qwen3-ASR-0.6BModel Scale=0.6B2026.05 | 2.81 | — | — | — | |
| Ark-Base+OPDModel Scale=0.6B, OPD=On-policy distillation2026.05 | 2.88 | — | — | — | |
| Gemini-2.5-Pro2026.01 | 2.89 | — | — | — | |
| Doubao-ASR-26032026.05 | 2.94 | — | — | — | |
| AR Greedy2026.06 | 3 | — | — | — | |
| NAR-MBR|Z|=64, Niter=12026.06 | 3.1 | — | — | — | |
| NAR-MBR|Z|=64, Niter=102026.06 | 3.1 | — | — | — | |
| NAR-MBR|Z|=256, Niter=12026.06 | 3.1 | — | — | — | |
| NAR-MBR|Z|=256, Niter=102026.06 | 3.1 | — | — | — | |
| Step-Audio-chat-3BParameters=3B, Note=Directly use official reported results2026.02 | 3.11 | — | — | — | |
| SLAM-CTCTrain Data (Audio)=Libri2026.04 | 3.13 | — | — | — | |
| NAR-MBR|Z|=256, Niter=02026.06 | 3.2 | — | — | — | |
| SeamlessM4T-v22026.01 | 3.3 | — | — | — | |
| SLAMTrain audio duration(h)=1.8k2026.04 | 3.3 | — | — | — | |
| NARNiter=02026.06 | 3.3 | — | — | — | |
| NARNiter=102026.06 | 3.3 | — | — | — | |
| NAR-MBR|Z|=64, Niter=02026.06 | 3.3 | — | — | — | |
| T-T2023.10 | 3.4 | 610 | — | — | |
| NARNiter=12026.06 | 3.4 | — | — | — | |
| TASU2Train Data (Text)=Libri, Conditioning=WER-binned2026.04 | 3.41 | — | — | — | |
| LLaMA-Omni22026.01 | 3.5 | — | — | — | |
| Mimo-Audio-Instruct-7BParameters=7B2026.02 | 3.5 | — | — | — | |
| Offline2023.10 | 3.51 | — | — | — | |
| OSUM-Pangu2026.03 | 3.51 | — | 3.51 | — | |
| Seg2Seg2023.10 | 3.55 | 324 | — | — | |
| TASU2Train audio duration(h)=02026.04 | 3.68 | — | — | — | |
| Whisper + Llama 2 + MMS TTSShots=10, Max context length=1024, Approach=Cascade Topline2024.02 | 3.7 | — | — | — | |
| Ark-BaseModel Scale=0.6B2026.05 | 3.75 | — | — | — | |
| MoChAConfig=DeCoT2023.10 | 3.9 | 240 | — | — | |
| Qwen2.5-Omni-7BParameters=7B2026.02 | 3.92 | — | — | — | |
| TASU2Train Data (Text)=Libri+Slide, Conditioning=WER-binned2026.04 | 3.94 | — | — | — | |
| T-TAlignment=ConstAlign2023.10 | 4 | 328 | — | — | |
| T-TAlignment=FastEmit2023.10 | 4 | 195 | — | — | |
| T-TAlignment=SelfAlign2023.10 | 4 | 145 | — | — | |
| MoChAConfig=CTC2023.10 | 4 | 240 | — | — |