Speech Quality Evaluation on OpenAudioBench English subsets (test)
2.31WERQwen2.5-Omni
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-Omni2026.02 | 2.31 | 4.34 | |
| VocalNet-8B2026.02 | 3.64 | 4.49 | |
| VITA-Audio2026.02 | 3.83 | 4.26 | |
| Baseline-ARPrediction Type=MTP2026.02 | 4.05 | 4.49 | |
| VocalNet-MDMPrediction Type=MDM, Diffusion Steps=42026.02 | 5.34 | 4.49 | |
| VocalNet-MDMPrediction Type=MDM, Diffusion Steps=162026.02 | 5.53 | 4.49 | |
| VocalNet-MDMPrediction Type=MDM, Diffusion Steps=82026.02 | 5.55 | 4.49 | |
| SLAM-Omni2026.02 | 5.78 | 4.46 | |
| VocalNet-MDMPrediction Type=MDM, Diffusion Steps=22026.02 | 6.1 | 4.47 | |
| MiMo-Audio2026.02 | 6.13 | 3.68 | |
| VocalNet-MDMPrediction Type=MDM, Diffusion Steps=12026.02 | 6.23 | 4.46 | |
| MiniCPM-o2026.02 | 9.52 | 4.14 | |
| Baseline-ARPrediction Type=NTP2026.02 | 10.66 | 4.48 | |
| GLM-4-Voice2026.02 | 11.9 | 4.23 | |
| Kimi-Audio2026.02 | 14.71 | 2.87 |