Speech Generation and Interaction on Medical Speech Dialogue Benchmark
3.96UTMOSZhongjing
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ZhongjingInput=Text/Audio, Output=Text2026.01 | 3.96 | 6.77 | 3,520 | |
| Qwen2-AudioInput=Text/Audio, Output=Text2026.01 | 3.96 | 11.83 | 4,072 | |
| SpeechMedAssistInput=Audio, Output=Audio2026.01 | 3.75 | 7.71 | 367 | |
| LLaMA-Omni2Input=Audio, Output=Audio2026.01 | 3.69 | 8.06 | 374 | |
| GLM4-VoiceInput=Audio, Output=Audio2026.01 | 3 | 15.3 | 1,562 | |
| Kimi-AudioInput=Audio, Output=Audio, Streaming Supported=false2026.01 | 2.55 | 4.94 | 3,134 | |
| SpeechGPT2Input=Audio, Output=Audio, Streaming Supported=false2026.01 | 2.49 | 15.3 | 8,470 |