Emotion Reasoning (ER) on AICA-Bench 1.0 (test)
0.4MSEAICA-Bench Scoring Model
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AICA-Bench Scoring Modelbase_model=QwenVL2.5-7B, training_regime=fine-tuned with human-labeled data2026.04 | 0.4 | 0.295 | 0.88 | |
| Qwen2.5VL-7B2026.04 | 1.22 | 0.775 | 0.472 | |
| ChatGPT-4o2026.04 | 1.36 | 0.79 | 0.502 | |
| Gemini2.5-Pro2026.04 | 1.94 | 0.94 | 0.441 |