Commonsense Reasoning on Spoken StoryCloze
89.1AccuracyASR + Qwen2.5-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ASR + Qwen2.5-7BSystem Category=Cascaded Toplines, ASR=Whisper-v3-Large, Backbone LLM=Qwen2.5-7B2026.05 | 89.1 | 0.1 | |
| WhisperPro + Qwen2.5-7BSystem Category=Ablation Studies, Backbone LLM=Qwen2.5-7B2026.05 | 88.8 | 0.4 | |
| TextPro-SLM-7BSystem Category=Modality Gap-optimized Systems2026.05 | 88.6 | 0.6 | |
| ASR + Qwen2.5-3BSystem Category=Cascaded Toplines, ASR=Whisper-v3-Large, Backbone LLM=Qwen2.5-3B2026.05 | 88.3 | 4.3 | |
| TextPro-SLM-7B 5:1System Category=Modality Gap-optimized Systems, Sampling Ratio=5:12026.05 | 87.6 | 1.6 | |
| TextPro-SLM-7B w/o KDSystem Category=Ablation Studies, Ablation Setting=w/o KD2026.05 | 84.7 | 4.5 | |
| TextPro-SLM-3BSystem Category=Modality Gap-optimized Systems2026.05 | 84.2 | 8.4 | |
| Qwen2.5-Omni-7BSystem Category=End-to-end Systems, Backbone LLM=Qwen2.5-Omni-7B2026.05 | 83.9 | 5.4 | |
| DiVA-Llama3.1-8BSystem Category=End-to-end Systems, Backbone LLM=Llama3.1-8B2026.05 | 82.1 | 15.8 | |
| SALAD-7BSystem Category=Modality Gap-optimized Systems, Training Stage=Stage I2026.05 | 81.5 | 3.5 | |
| SALAD-7BSystem Category=Modality Gap-optimized Systems, Training Stage=Stage II2026.05 | 81.5 | 3.5 | |
| GLM-4-Voice-9BSystem Category=End-to-end Systems, Backbone LLM=GLM-4-Voice-9B2026.05 | 76.4 | 20.6 | |
| SALAD-3BSystem Category=Modality Gap-optimized Systems, Training Stage=Stage II2026.05 | 75.8 | 7.1 | |
| SALAD-3BSystem Category=Modality Gap-optimized Systems, Training Stage=Stage I2026.05 | 75.5 | 7.4 | |
| Qwen2-Audio-7BSystem Category=End-to-end Systems, Backbone LLM=Qwen2-Audio-7B2026.05 | 71.9 | 9 | |
| Kimi-Audio-7BSystem Category=End-to-end Systems, Backbone LLM=Kimi-Audio-7B2026.05 | 66.6 | 22.6 | |
| Random2026.05 | 50 | — |