Spoken Question Answering on Our Bench
76.34AccuracySerial
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Serialthinking_mode=Think, variant=upper bound2026.01 | 76.34 | 51,000 | 61,000 | |
| LTS-VADthinking_mode=Think2026.01 | 72.36 | 29,000 | 38,000 | |
| PredGenthinking_mode=Think2026.01 | 70.51 | 43,000 | 48,000 | |
| LTS-VoiceAgent2026.01 | 62.57 | 236 | 336 | |
| Serialthinking_mode=No-think2026.01 | 58.45 | 663 | 1,764 | |
| LTS-VADthinking_mode=No-think2026.01 | 56.41 | 244 | 294 | |
| PredGenthinking_mode=No-think2026.01 | 55.84 | 251 | 303 | |
| Qwen2.5-Omniarchitecture=E2E2026.01 | 50.25 | 245.17 | 449.9 |