Automatic Speech Recognition on SlideASR S (en) 1.0
4.6WERVAPO-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VAPO-7BContext Setting=Slide image as context (End-to-End)2025.10 | 4.6 | 2.83 | 2.97 | |
| VAPO-3BContext Setting=Slide image as context (End-to-End)2025.10 | 4.9 | 3.19 | 3.73 | |
| Qwen2.5-Omni-7BContext Setting=Contextless2025.10 | 8.15 | 23.44 | 27.77 | |
| Qwen2.5-Omni-3BContext Setting=Contextless2025.10 | 8.37 | 24.15 | 31.04 | |
| Qwen3-Omni-30B-A3BContext Setting=Contextless2025.10 | 9.06 | 14.61 | 15.53 | |
| MiniCPM-o-2.6Context Setting=Contextless2025.10 | 11.19 | 27.51 | 30.93 | |
| Qwen2-AudioContext Setting=Contextless2025.10 | 11.9 | 36.29 | 47.84 | |
| Qwen3-Omni-30B-A3BContext Setting=Slide text as context (Pipeline)2025.10 | 34.65 | 32.35 | 8.56 | |
| Qwen2.5-Omni-7BContext Setting=Slide image as context (End-to-End)2025.10 | 57.21 | 35.76 | 15.04 | |
| Mi-DashengContext Setting=Slide text as context (Pipeline)2025.10 | 78.98 | 49.85 | 30.58 | |
| Qwen2-AudioContext Setting=Slide text as context (Pipeline)2025.10 | 92.16 | 66.38 | 24.82 | |
| Qwen2.5-Omni-3BContext Setting=Slide image as context (End-to-End)2025.10 | 100.08 | 53.19 | 18.72 | |
| Qwen3-Omni-30B-A3BContext Setting=Slide image as context (End-to-End)2025.10 | 101.45 | 59.64 | 12.08 | |
| MiniCPM-o-2.6Context Setting=Slide image as context (End-to-End)2025.10 | 112.9 | 49.65 | 15.01 |