Automatic Speech Recognition on Librispeech
0.0141WERKimi-Audio
Evaluation Results
| Method | Links | |
|---|---|---|
| Kimi-AudioModel Category=Medium Sized Large Audio Language Models, Parameter Count=5B-20B parameters2025.09 | 0.0141 | |
| Voxtral-Small-24BModel Category=Large Sized Large Audio Language Models, Parameter Count=>20B parameters2025.09 | 0.0162 | |
| Qwen3-Omni-30B-A3B-ThinkingModel Category=Large Sized Large Audio Language Models, Parameter Count=>20B parameters2025.09 | 0.0164 | |
| Qwen2.5-Omni-7BModel Category=Medium Sized Large Audio Language Models, Parameter Count=5B-20B parameters2025.09 | 0.0174 | |
| Phi-4-Multi-modalModel Category=Medium Sized Large Audio Language Models, Parameter Count=5B-20B parameters2025.09 | 0.0197 | |
| Polyglot-Lion-1.7BParams=1.7B2026.03 | 0.021 | |
| Voxtral-Mini-3BModel Category=Small-sized Audio Language Models, Parameter Count=<5B parameters2025.09 | 0.021 | |
| Gemini2.5-FlashModel Category=Proprietary Audio Language Models2025.09 | 0.0217 | |
| Qwen3-ASR-1.7BParams=1.7B2026.03 | 0.0231 | |
| MERaLiON-2-10B-ASRParams=10B2026.03 | 0.0254 | |
| Polyglot-Lion-0.6BParams=0.6B2026.03 | 0.0267 | |
| Qwen3-ASR-0.6BParams=0.6B2026.03 | 0.0274 | |
| GLM-4-VoiceModel Type=Speech LLM2026.03 | 0.0282 | |
| EMOVAModel Size=72B2024.09 | 0.029 | |
| WhisperModel Size=Large2024.09 | 0.03 | |
| Whisper-large-v3-turboParams=0.8B2026.03 | 0.0304 | |
| VITAModel Size=8x7B2024.09 | 0.034 | |
| EMOVAModel Size=7B2024.09 | 0.041 | |
| GPT-4o-transcribe + GPT-4.1-miniModel Category=Cascaded Systems2025.09 | 0.0471 | |
| Mini-Omni22024.09 | 0.048 | |
| Whisper (No SAE)Steering Method=None, SAE=None2026.02 | 0.051 | |
| Whisper + AudioSAESteering Method=None, Configuration=Injected SAE on last layer2026.02 | 0.052 | |
| S-Vector SteeringAlpha (α)=3, Calculation Dataset=Musan2026.02 | 0.053 | |
| EMOVAModel Size=3B2024.09 | 0.054 | |
| AudioSAE SteeringAlpha (α)=1, Features=Top-100 from FSD50k2026.02 | 0.055 | |
| GPT-4o-mini-audioModel Category=Proprietary Audio Language Models2025.09 | 0.0625 | |
| Omni-DiffusionModel Type=Any-to-Any2026.03 | 0.0705 | |
| Qwen2.5-Omni-3BModel Category=Small-sized Audio Language Models, Parameter Count=<5B parameters2025.09 | 0.0809 | |
| VITAModel Size=1.52024.09 | 0.081 | |
| AnyGPTModel Type=Any-to-Any2026.03 | 0.085 | |
| Whisper-Large-v3 + GPT-oSS-20BModel Category=Cascaded Systems2025.09 | 0.0982 | |
| Qwen2.5-Omni-7BParams=7B2026.03 | 0.138 | |
| Qwen2.5-Omni-3BParams=3B2026.03 | 0.2921 | |
| SeaLLMs-Audio-7BParams=7B2026.03 | 0.9474 | |
| AudioSAE SteeringAlpha (α)=3, Features=Top-100 from FSD50k2026.02 | 0.984 |