Multi-modal Stance Detection on MTWQ
65MOC Macro F1GPT-4 Vision
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4 VisionModality=Multi-modal, Zero-shot=true2026.04 | 65 | 52.36 | |
| MM-StanceDetModality=Multi-modal, Zero-shot=true2026.04 | 63.02 | 51.4 | |
| MV-DebateModality=Textual, Zero-shot=true2026.04 | 62.66 | 52.18 | |
| GPT-4 + CoTModality=Textual, Zero-shot=true2026.04 | 62.4 | 52.41 | |
| GPT-4Modality=Textual, Zero-shot=true2026.04 | 62.1 | 52.12 | |
| BridgeTowerModality=Multi-modal, Zero-shot=true2026.04 | 61.59 | 49.72 | |
| LLaMA2Modality=Textual, Zero-shot=true2026.04 | 51.46 | 44.1 | |
| Qwen-VLModality=Multi-modal, Zero-shot=true2026.04 | 44.32 | 44.08 | |
| TASTEModality=Multi-modal, Zero-shot=true2026.04 | 42.19 | 40.88 | |
| TMPTModality=Multi-modal, Zero-shot=true2026.04 | 32.18 | 26.48 | |
| RoBERTaModality=Textual, Zero-shot=true2026.04 | 30.62 | 15.84 | |
| LKI-BARTModality=Textual, Zero-shot=true2026.04 | 29.54 | 20.16 | |
| ViTModality=Visual, Zero-shot=true2026.04 | 29.37 | 23.69 | |
| KEBERTModality=Textual, Zero-shot=true2026.04 | 29.17 | 19.8 | |
| BERTModality=Textual, Zero-shot=true2026.04 | 28.04 | 9.57 | |
| SwinTModality=Visual, Zero-shot=true2026.04 | 27.9 | 19.69 | |
| ResNetModality=Visual, Zero-shot=true2026.04 | 27.59 | 24.88 | |
| CLIPModality=Multi-modal, Zero-shot=true2026.04 | 27.21 | 15.69 | |
| BERT+ViTModality=Multi-modal, Zero-shot=true2026.04 | 24.76 | 11.7 | |
| ViLTModality=Multi-modal, Zero-shot=true2026.04 | 23.54 | 19.18 |