Commonsense Reasoning on PIQA (Accuracy (%), Modality Gap)
82.9AccuracyKimi-Audio-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Kimi-Audio-7BSystem Category=End-to-end Systems, Backbone LLM=Kimi-Audio-7B2026.05 | 82.9 | -14.8 | |
| SALAD-7BSystem Category=Modality Gap-optimized Systems, Training Stage=Stage I2026.05 | 80.3 | 0.4 | |
| SALAD-7BSystem Category=Modality Gap-optimized Systems, Training Stage=Stage II2026.05 | 80.3 | 0.4 | |
| SALAD-3BSystem Category=Modality Gap-optimized Systems, Training Stage=Stage I2026.05 | 78.3 | 0.3 | |
| SALAD-3BSystem Category=Modality Gap-optimized Systems, Training Stage=Stage II2026.05 | 78.1 | 0.5 | |
| Qwen2-Audio-7BSystem Category=End-to-end Systems, Backbone LLM=Qwen2-Audio-7B2026.05 | 73.4 | 5.4 | |
| Qwen2.5-Omni-7BSystem Category=End-to-end Systems, Backbone LLM=Qwen2.5-Omni-7B2026.05 | 72.8 | -4.8 | |
| TextPro-SLM-7B 5:1System Category=Modality Gap-optimized Systems, Sampling Ratio=5:12026.05 | 70.8 | -2.7 | |
| TextPro-SLM-7BSystem Category=Modality Gap-optimized Systems2026.05 | 69.7 | -1.6 | |
| TextPro-SLM-7B w/o KDSystem Category=Ablation Studies, Ablation Setting=w/o KD2026.05 | 66.6 | 1.5 | |
| WhisperPro + Qwen2.5-7BSystem Category=Ablation Studies, Backbone LLM=Qwen2.5-7B2026.05 | 63.3 | 4.8 | |
| ASR + Qwen2.5-7BSystem Category=Cascaded Toplines, ASR=Whisper-v3-Large, Backbone LLM=Qwen2.5-7B2026.05 | 63 | 5.1 | |
| TextPro-SLM-3BSystem Category=Modality Gap-optimized Systems2026.05 | 56.9 | -4.8 | |
| ASR + Qwen2.5-3BSystem Category=Cascaded Toplines, ASR=Whisper-v3-Large, Backbone LLM=Qwen2.5-3B2026.05 | 52.7 | -0.6 | |
| Random2026.05 | 50 | — | |
| GLM-4-Voice-9BSystem Category=End-to-end Systems, Backbone LLM=GLM-4-Voice-9B2026.05 | 47.3 | 30.9 | |
| DiVA-Llama3.1-8BSystem Category=End-to-end Systems, Backbone LLM=Llama3.1-8B2026.05 | 35.6 | 30.9 |