Automatic Speech Recognition on SlideASR 1.0 (R)
26.48NE-WERVAPO-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VAPO-7BContext Setting=Slide image as context (End-to-End)2025.10 | 26.48 | 15.35 | |
| VAPO-3BContext Setting=Slide image as context (End-to-End)2025.10 | 27.28 | 19.31 | |
| Qwen3-Omni-30B-A3BContext Setting=Slide image as context (End-to-End)2025.10 | 32.26 | 24.75 | |
| Qwen3-Omni-30B-A3BContext Setting=Slide text as context (Pipeline)2025.10 | 34.01 | 28.22 | |
| Qwen3-Omni-30B-A3BContext Setting=Contextless2025.10 | 40.43 | 41.09 | |
| Qwen2.5-Omni-7BContext Setting=Slide image as context (End-to-End)2025.10 | 41.77 | 35.15 | |
| Mi-DashengContext Setting=Slide text as context (Pipeline)2025.10 | 47.52 | 26.73 | |
| Qwen2.5-Omni-3BContext Setting=Slide image as context (End-to-End)2025.10 | 49 | 53.47 | |
| Qwen2.5-Omni-7BContext Setting=Contextless2025.10 | 53.68 | 63.37 | |
| MiniCPM-o-2.6Context Setting=Contextless2025.10 | 55.85 | 65.37 | |
| Qwen2-AudioContext Setting=Slide text as context (Pipeline)2025.10 | 59.04 | 21.29 | |
| Qwen2.5-Omni-3BContext Setting=Contextless2025.10 | 61.31 | 66.83 | |
| MiniCPM-o-2.6Context Setting=Slide image as context (End-to-End)2025.10 | 63.73 | 66.83 | |
| Qwen2-AudioContext Setting=Contextless2025.10 | 74.56 | 76.73 |