Multimodal Emotion Reasoning on EMER
7.83Clue OverlapEmotion-LLaMA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Emotion-LLaMA2024.06 | 7.83 | 6.25 | |
| Emotion-LLaMAv22026.01 | 7.3 | 7.14 | |
| Valley2024.06 | 7.24 | 5.77 | |
| VideoChat-EmbedInput format=textual embedding space2024.06 | 7.15 | 5.65 | |
| PandaGPT2024.06 | 7.14 | 5.51 | |
| Video-ChatGPT2024.06 | 6.95 | 5.74 | |
| Video-LLaMA2024.06 | 6.64 | 4.89 | |
| VideoChat-TextInput format=textual format2024.06 | 6.42 | 3.94 | |
| Emotion-LLaMA2026.01 | 5.89 | 6.89 | |
| AffectGPT2026.01 | 5.87 | 5.79 | |
| Video-LLaMA22026.01 | 5.63 | 5.51 | |
| Qwen2.5-Omni2026.01 | 5.6 | 5.04 | |
| PandaGPT2026.01 | 4.86 | 4.39 | |
| Video-ChatGPT2026.01 | 4.81 | 5.07 | |
| Valley2026.01 | 4.32 | 5.45 | |
| VideoChatvariant=Embed2026.01 | 4.04 | 5.12 | |
| Video-LLaMA2026.01 | 3.68 | 5.01 | |
| VideoChatvariant=Text2026.01 | 3.28 | 3.89 |