Automatic Speech Recognition on SlideASR S (zh) 1.0
2.13WERVAPO-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VAPO-7BContext Setting=Slide image as context (End-to-End)2025.10 | 2.13 | 3.78 | 1.36 | |
| VAPO-3BContext Setting=Slide image as context (End-to-End)2025.10 | 2.47 | 4.21 | 2.22 | |
| Qwen2.5-Omni-7BContext Setting=Contextless2025.10 | 4.34 | 17.54 | 32.8 | |
| Qwen2.5-Omni-3BContext Setting=Contextless2025.10 | 4.47 | 19.89 | 38.08 | |
| Qwen2-AudioContext Setting=Contextless2025.10 | 6.02 | 22.83 | 40.36 | |
| Qwen3-Omni-30B-A3BContext Setting=Slide text as context (Pipeline)2025.10 | 9.76 | 15.85 | 13.54 | |
| MiniCPM-o-2.6Context Setting=Contextless2025.10 | 10.35 | 25 | 41.62 | |
| Qwen3-Omni-30B-A3BContext Setting=Contextless2025.10 | 20.77 | 23.31 | 22.49 | |
| Qwen2-AudioContext Setting=Slide text as context (Pipeline)2025.10 | 39.09 | 50.58 | 31.52 | |
| Mi-DashengContext Setting=Slide text as context (Pipeline)2025.10 | 66.88 | 56.3 | 32.3 | |
| Qwen3-Omni-30B-A3BContext Setting=Slide image as context (End-to-End)2025.10 | 79.21 | 46.45 | 5.54 | |
| Qwen2.5-Omni-3BContext Setting=Slide image as context (End-to-End)2025.10 | 86.86 | 65.62 | 9.62 | |
| MiniCPM-o-2.6Context Setting=Slide image as context (End-to-End)2025.10 | 89.53 | 61.25 | 45.67 | |
| Qwen2.5-Omni-7BContext Setting=Slide image as context (End-to-End)2025.10 | 91.83 | 54.04 | 3.36 |