Harmful Meme Detection on FHM
75.79Macro-F1ALARM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ALARMTraining Free=true, Setting=Label-Free, Backbone=Qwen2.5-VL-72B2025.12 | 75.79 | 75.8 | |
| ExplainHMTraining Free=false, Setting=Label-Driven2025.12 | 75.39 | 75.6 | |
| MR.HARMTraining Free=false, Setting=Label-Driven2025.12 | 75.1 | 75.4 | |
| ALARMTraining Free=true, Setting=Label-Free, Backbone=GPT-4o2025.12 | 72.75 | 72.8 | |
| PromptHateTraining Free=false, Setting=Label-Driven2025.12 | 71.83 | 72.2 | |
| Pro-CapTraining Free=false, Setting=Label-Driven2025.12 | 71.68 | 74.95 | |
| Qwen2.5-VL-72BTraining Free=true, Setting=Few Shot2025.12 | 71.02 | 71.2 | |
| LoReHMTraining Free=true, Setting=Few Shot, Backbone=GPT-4o2025.12 | 70.14 | 70.2 | |
| HHPromptTraining Free=false, Setting=Label-Driven2025.12 | 69.01 | 70.4 | |
| ISMTraining Free=false, Setting=Label-Driven2025.12 | 68.77 | 70.45 | |
| LoReHMTraining Free=true, Setting=Few Shot, Backbone=Qwen2.5-VL-72B2025.12 | 68.67 | 69 | |
| GPT-4oSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 68.25 | 68.8 | |
| GPT-4oProtocol=Zero-Shot, Model Category=Closed-Source VLMs2026.05 | 68.25 | 68.8 | |
| PrismAgentProtocol=Zero-Shot, Model Category=Agent-Based Methods (Open-Source), Backbone=LLaVA-1.6-34B2026.05 | 66.72 | 66.8 | |
| GPT-4oTraining Free=true, Setting=Few Shot2025.12 | 65.74 | 66.6 | |
| Label-tunedSource Dataset=MuPHI2026.05 | 64.4 | — | |
| PrismAgentProtocol=Zero-Shot, Model Category=Agent-Based Methods (Open-Source), Backbone=LLaVA-1.5-13B2026.05 | 63.96 | 64 | |
| LLaVA-1.6-34BSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 63.51 | 64 | |
| LLaVA-1.6-34BProtocol=Zero-Shot, Model Category=Open-Source VLMs2026.05 | 63.51 | 64 | |
| MuPHIRMSource Dataset=MuPHI2026.05 | 62.7 | — | |
| MuPHIRMSource Dataset=Harm-P2026.05 | 61.8 | — | |
| MuPHIRMSource Dataset=Harm-C2026.05 | 61.2 | — | |
| MINDBackbone=LLaVA-1.5-13B, Training/Evaluation Protocol=Training-free2025.07 | 60.71 | 60.8 | |
| MINDBackbone=LLaVA-1.5-13B, Setting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 60.71 | 60.8 | |
| MINDProtocol=Zero-Shot, Model Category=Agent-Based Methods (Open-Source), Backbone=LLaVA-1.5-13B2026.05 | 60.71 | 60.8 | |
| Label-tunedSource Dataset=Harm-P2026.05 | 59.1 | — | |
| Gemini 1.5 FlashSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 58.9 | 60.2 | |
| MOMENTATraining/Evaluation Protocol=Training-based2025.07 | 57.45 | 61.34 | |
| MOMENTATraining Free=false, Setting=Label-Driven2025.12 | 57.45 | 61.34 | |
| MOMENTAProtocol=Supervised / Training-Based2026.05 | 57.45 | 61.34 | |
| Gemini-2.0-FlashProtocol=Zero-Shot, Model Category=Closed-Source VLMs2026.05 | 54.04 | 60.4 | |
| Mod-HATETraining Free=false, Setting=Few Shot2025.12 | 53.88 | 57.6 | |
| LLaVA-1.5-13BSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 53.01 | 55.2 | |
| LLaVA-1.5-13BProtocol=Zero-Shot, Model Category=Open-Source VLMs2026.05 | 53.01 | 55.2 | |
| InstructBLIP-13BSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 51.89 | 55.4 | |
| InstructBLIP-13BProtocol=Zero-Shot, Model Category=Open-Source VLMs2026.05 | 51.89 | 55.4 | |
| OpenFlamingo-9BTraining Free=true, Setting=Few Shot2025.12 | 51.52 | 51.6 | |
| OPT-30BTraining Free=true, Setting=Few Shot2025.12 | 50.82 | 54.2 | |
| OpenFlamingo-9BSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 49.52 | 50.5 | |
| OpenFlamingo-9BProtocol=Zero-Shot, Model Category=Open-Source VLMs2026.05 | 49.52 | 50.5 | |
| InstructBLIP-7BSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 48.85 | 52 | |
| InstructBLIP-7BProtocol=Zero-Shot, Model Category=Open-Source VLMs2026.05 | 48.85 | 52 | |
| MiniGPT-v2-7BSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 47.88 | 51.3 | |
| MiniGPT-V-2BProtocol=Zero-Shot, Model Category=Open-Source VLMs2026.05 | 47.88 | 51.3 | |
| LLaVA-1.5-7BSetting=Zero-shot, Prompt Strategy=Chain-of-Thought2025.07 | 45.51 | 53.8 | |
| LLaVA-1.5-7BProtocol=Zero-Shot, Model Category=Open-Source VLMs2026.05 | 45.51 | 53.8 | |
| Late FusionTraining/Evaluation Protocol=Training-based2025.07 | 44.81 | 59.14 | |
| Late FusionProtocol=Supervised / Training-Based2026.05 | 44.81 | 59.14 | |
| Label-tunedSource Dataset=Harm-C2026.05 | 33.3 | — |