Dimensional Emotion Recognition on IEMOCAP
231Count VVQwen-Omni
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Qwen-OmniBackbone=Streaming-modified Whisper, LLM=Qwen2.5 Thinker, Adapter=Learned projection2025.05 | 231 | 64 | 51 | 27 | 1 | 0.99 | 0.433 | 0.436 | |
| AF3Backbone=AF-Whisper, LLM=Qwen2.5-7B2025.05 | 222 | 73 | 44 | 28 | 1.03 | 1.04 | 0.383 | 0.378 | |
| Qwen2-AudioBackbone=Whisper-large-v3, LLM=Qwen-7B, Adapter=None2025.05 | 190 | 59 | 46 | 16 | 1.05 | 1.05 | 0.396 | 0.398 | |
| SALMONNBackbone=Whisper-v2 + BEATs, LLM=Vicuna-13B, Adapter=Q-Former2025.05 | 183 | 50 | 35 | 18 | 1.04 | 1.1 | 0.398 | 0.367 |