Automatic Speech Recognition on LibriSpeech (other)
2.42WERKimi Audio
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Kimi Audio2026.02 | 2.42 | — | — | |
| Kimi-Audio2025.12 | 2.42 | — | — | |
| Samba-ASRorganization=SandLogic2025.01 | 2.48 | — | — | |
| Qwen3-Omni-Instruct2026.02 | 2.48 | — | — | |
| nvidia/parakeet-tdt-1.1bparameters=1.1b2025.01 | 2.6 | — | — | |
| ERNIE 5.02026.02 | 2.61 | — | — | |
| Kimi Audio2026.04 | 2.7 | — | — | |
| Step-Audio2-Mini2025.12 | 2.86 | — | — | |
| Raon-Speech2026.04 | 2.89 | — | — | |
| nvidia/canary-1bparameters=1b2025.01 | 2.93 | — | — | |
| Step-Audio Flamingo 32026.04 | 2.97 | — | — | |
| Audio-FlamingoModel version=32025.12 | 3.13 | — | — | |
| Audio Flamingo 3Backbone=Audio Flamingo 3, Method=Baseline2026.06 | 3.13 | — | — | |
| StepAudio 2.5 ASRMTP training=false2026.05 | 3.14 | — | — | |
| StepAudio 2.5 ASRMTP training=true2026.05 | 3.16 | — | — | |
| Qwen3-ASR-1.7BParameters=1.7B2026.01 | 3.38 | — | — | |
| Qwen2.5-Omni-7BBackbone=Qwen2.5-Omni-7B, Method=Baseline2026.06 | 3.4 | — | — | |
| Qwen2-AudioType=AR2026.01 | 3.5 | — | — | |
| Phi-4 Multimodal2026.01 | 3.5 | — | — | |
| Whisper-large-v3Model size=large2026.02 | 3.55 | — | — | |
| Gemini-2.5-Pro2026.01 | 3.56 | — | — | |
| MiniCPM-o 4.52026.04 | 3.56 | — | — | |
| Qwen3-ASR-1.7B2026.05 | 3.57 | — | — | |
| Phi-4-MultimodalType=AR2026.01 | 3.6 | — | — | |
| Qwen2-Audio2026.01 | 3.6 | — | — | |
| Qwen2-Audio2026.03 | 3.6 | — | 3.6 | |
| Mixed Training + Length AugLength Augmentation=true2026.04 | 3.65 | — | — | |
| Mixed Training + Length Aug + Reduced TFLength Augmentation=true, Reduced Teacher Forcing=true2026.04 | 3.65 | — | — | |
| ASR-only Baseline2026.04 | 3.68 | — | — | |
| Mixed Training + Length Aug + Timestamp RegLength Augmentation=true, Timestamp Regularization=true2026.04 | 3.69 | — | — | |
| MoST2026.01 | 3.7 | — | — | |
| CrisperWhisper2026.04 | 3.72 | — | — | |
| ERNIE 5.0-Base2026.02 | 3.73 | — | — | |
| Fun-Audio-ChatModel size=30B-A3B2025.12 | 3.73 | — | — | |
| GPT-4o-Audio2026.02 | 3.75 | — | — | |
| GPT-4o-Transcribe2026.01 | 3.75 | — | — | |
| DAPlin=✓, ldeep=✓, FRR=33.76, Pre. GFLOPs=686.39, FLOPs Ratio=87.89, Setting=Conservative (τin=0.90, τdeep=0.80)2026.04 | 3.77 | — | — | |
| DAPlin=✓, ldeep=✓, FRR=14.91, Pre. GFLOPs=566.30, FLOPs Ratio=72.52, Setting=Aggressive (τin=0.80, τdeep=0.70)2026.04 | 3.79 | — | — | |
| APinlin=✓, ldeep=-, FRR=78.64, Pre. GFLOPs=612.93, FLOPs Ratio=78.49, Setting=Aggressive (τin=0.80, τdeep=0.70)2026.04 | 3.8 | — | — | |
| APinlin=✓, ldeep=-, FRR=93.56, Pre. GFLOPs=730.28, FLOPs Ratio=93.51, Setting=Conservative (τin=0.90, τdeep=0.80)2026.04 | 3.81 | — | — | |
| Mixed Training2026.04 | 3.82 | — | — | |
| APdeeplin=-, ldeep=✓, FRR=33.29, Pre. GFLOPs=731.96, FLOPs Ratio=93.73, Setting=Conservative (τin=0.90, τdeep=0.80)2026.04 | 3.83 | — | — | |
| APdeeplin=-, ldeep=✓, FRR=14.30, Pre. GFLOPs=718.12, FLOPs Ratio=91.96, Setting=Aggressive (τin=0.80, τdeep=0.70)2026.04 | 3.84 | — | — | |
| Vanillalin=-, ldeep=-, FRR=100.0, Pre. GFLOPs=780.94, FLOPs Ratio=100.0, Setting=N/A2026.04 | 3.88 | — | — | |
| Qwen2.5-Omni2026.04 | 3.88 | — | — | |
| Fun-Audio Chat2026.04 | 3.89 | — | — | |
| MinMoType=AR2026.01 | 3.9 | — | — | |
| MinMo2026.01 | 3.9 | — | — | |
| openai/whisper-large-v3model_size=large-v32025.01 | 3.91 | — | — | |
| Whisper-large-v32026.01 | 3.97 | — | — | |
| nyrahealth/CrisperWhisper2025.01 | 4 | — | — | |
| Llama-Omni2Type=AR2026.01 | 4 | — | — | |
| LLaMA-Omni22026.01 | 4 | — | — | |
| Qwen2.5-Omni-7B + CoATBackbone=Qwen2.5-Omni-7B, Method=CoAT2026.06 | 4 | — | — | |
| LongCat-Flash-Omni2026.02 | 4.01 | — | — | |
| FunASR-MLT-Nano2026.01 | 4.03 | — | — | |
| Qwen3-ASR-1.7BModel Scale=1.7B2026.05 | 4.05 | — | — | |
| Fun-Audio-ChatModel size=8B2025.12 | 4.13 | — | — | |
| SeamlessM4T-v22026.01 | 4.2 | — | — | |
| Qwen2.5-OmniTrain audio duration(h)=>1000k2026.04 | 4.21 | — | — | |
| Audio Flamingo 3 + CoATBackbone=Audio Flamingo 3, Method=CoAT2026.06 | 4.23 | — | — | |
| DIFFUSPEECHType=Diff.2026.01 | 4.3 | — | — | |
| Gemini-3-Pro2026.02 | 4.4 | — | — | |
| FunASR-Nano2026.05 | 4.43 | — | — | |
| Qwen3-ASR-0.6BParameters=0.6B2026.01 | 4.55 | — | — | |
| Ark-Base+TD+OPDModel Scale=0.6B, TD=Teacher-data adaptation, OPD=On-policy distillation2026.05 | 4.56 | — | — | |
| Qwen-Audio2026.04 | 4.59 | — | — | |
| Fun-Audio Omni2026.04 | 4.67 | — | — | |
| Qwen2-Audio + CoATBackbone=Qwen2-Audio, Method=CoAT2026.06 | 4.74 | — | — | |
| HyperCLOVA X 8B Omni2026.04 | 5.03 | — | — | |
| Qwen3-ASR-0.6BModel Scale=0.6B2026.05 | 5.05 | — | — | |
| Seamless ASRBackbone=Seamless2026.01 | 5.1 | — | 2 | |
| Qwen2.5-Omni-7Bevaluation_setting=Seen-task evaluation2026.02 | 5.5 | — | — | |
| Ark-Base+OPDModel Scale=0.6B, OPD=On-policy distillation2026.05 | 5.5 | — | — | |
| AR Beam2026.06 | 5.5 | — | — | |
| Qwen2.5-Omni-7BParameters=7B2026.02 | 5.52 | — | — | |
| OSUM2026.03 | 5.53 | — | 5.53 | |
| Doubao-ASR2026.01 | 5.7 | — | — | |
| VibeVoice-ASR2026.05 | 5.79 | — | — | |
| SpeechMapper (EuroLLM)Backbone=EuroLLM, Stage=stage 2, Objective=CE+MSE2026.01 | 5.8 | — | 2.7 | |
| SpeechMapper (Llama 3.1)Backbone=Llama 3.1, Stage=stage 2, Objective=ASR CE2026.01 | 5.8 | — | 2.7 | |
| Doubao-ASR-26032026.05 | 5.98 | — | — | |
| SpeechMapper (EuroLLM)Backbone=EuroLLM, Stage=stage 2, Objective=ASR CE2026.01 | 6 | — | 2.7 | |
| AR Greedy2026.06 | 6 | — | — | |
| UniAudio 2.0evaluation_setting=Seen-task evaluation2026.02 | 6.3 | — | — | |
| Step-AudioTrain audio duration(h)=>1000k2026.04 | 6.32 | — | — | |
| UniAudio 2.02026.02 | 6.33 | — | — | |
| Interactive 2 mini2026.04 | 6.82 | — | — | |
| NAR-MBR|Z|=256, Niter=12026.06 | 7 | — | — | |
| NAR-MBR|Z|=256, Niter=102026.06 | 7 | — | — | |
| Qwen2-AudioBackbone=Qwen2-Audio, Method=Baseline2026.06 | 7.02 | — | — | |
| NAR-MBR|Z|=64, Niter=12026.06 | 7.1 | — | — | |
| NAR-MBR|Z|=64, Niter=102026.06 | 7.1 | — | — | |
| Ark-BaseModel Scale=0.6B2026.05 | 7.17 | — | — | |
| Whisper + Llama 2 + MMS TTSShots=10, Max context length=1024, Approach=Cascade Topline2024.02 | 7.2 | — | — | |
| BEST-IWSLT25-IFBackbone=BEST-IWSLT25-IF2026.01 | 7.2 | — | 3.1 | |
| SLAMTrain audio duration(h)=1.8k2026.04 | 7.24 | — | — | |
| NARNiter=02026.06 | 7.4 | — | — | |
| NAR-MBR|Z|=64, Niter=02026.06 | 7.4 | — | — | |
| NAR-MBR|Z|=256, Niter=02026.06 | 7.4 | — | — |