Multi-modal Stance Detection on MRUC
45.37RUS Macro F1MM-StanceDet
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MM-StanceDetModality=Multi-modal, Zero-shot=true2026.04 | 45.37 | 55.25 | |
| GPT-4 VisionModality=Multi-modal, Zero-shot=true2026.04 | 42.09 | 47 | |
| MV-DebateModality=Textual, Zero-shot=true2026.04 | 40.81 | 49.23 | |
| GPT-4 + CoTModality=Textual, Zero-shot=true2026.04 | 40.55 | 49.45 | |
| GPT-4Modality=Textual, Zero-shot=true2026.04 | 40.22 | 49.18 | |
| BridgeTowerModality=Multi-modal, Zero-shot=true2026.04 | 39.85 | 45.33 | |
| Qwen-VLModality=Multi-modal, Zero-shot=true2026.04 | 36.95 | 41.39 | |
| TASTEModality=Multi-modal, Zero-shot=true2026.04 | 35.11 | 37.45 | |
| LLaMA2Modality=Textual, Zero-shot=true2026.04 | 31.86 | 36.34 | |
| ViTModality=Visual, Zero-shot=true2026.04 | 27.26 | 28.51 | |
| RoBERTaModality=Textual, Zero-shot=true2026.04 | 27.1 | 19.98 | |
| CLIPModality=Multi-modal, Zero-shot=true2026.04 | 25.62 | 27.4 | |
| SwinTModality=Visual, Zero-shot=true2026.04 | 25.44 | 24.54 | |
| LKI-BARTModality=Textual, Zero-shot=true2026.04 | 24.92 | 28.43 | |
| KEBERTModality=Textual, Zero-shot=true2026.04 | 24.68 | 28.18 | |
| ResNetModality=Visual, Zero-shot=true2026.04 | 23.88 | 25.57 | |
| TMPTModality=Multi-modal, Zero-shot=true2026.04 | 23.87 | 24.71 | |
| BERT+ViTModality=Multi-modal, Zero-shot=true2026.04 | 23.33 | 15.21 | |
| BERTModality=Textual, Zero-shot=true2026.04 | 22.01 | 15.45 | |
| ViLTModality=Multi-modal, Zero-shot=true2026.04 | 21.56 | 23.96 |