Speech-to-Speech Translation on CVSS-T zh→en
23.6ASR-BLEUSeamlessM4T-Large-v2
Evaluation Results
| Method | Links | |
|---|---|---|
| SeamlessM4T-Large-v22026.04 | 23.6 | |
| MoVE2026.04 | 21.4 | |
| Kimi + LoRATraining data source=Ours, Training duration (hours)=100h2026.04 | 21.2 | |
| Kimi + LoRATraining data source=Ours, Training duration (hours)=50h2026.04 | 20.1 | |
| gpt-4o-audio-preview2026.04 | 19.2 | |
| Kimi + LoRATraining data source=SynStard, Training duration (hours)=100h2026.04 | 18.4 | |
| SeamlessExpressive2026.04 | 18.2 | |
| Kimi + LoRATraining data source=SeamlessAlign, Training duration (hours)=67h2026.04 | 12.5 | |
| Kimi-Audio-7B-Instruct2026.04 | 11.2 | |
| CascadedOne-shot acoustic prompt=true2026.04 | 10.6 |