Binary Manipulation Detection on MMFakeBench (1000 samples, val)
74.9F1 ScoreREFORM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| REFORMLanguage Model=Florence2-0.3B, Prompt Method=Ours (Sec.3.2)2026.03 | 74.9 | 74.5 | 75.4 | 64.7 | |
| REFORMLanguage Model=Florence2-0.3B, Prompt=Ours (Sec.3.2), Zero-shot=true2026.03 | 74.9 | 74.5 | 75.4 | 64.7 | |
| VILALanguage Model=LLaMA2-13B, Prompt=MMD-Agent, Zero-shot=true2026.03 | 56.5 | 62.2 | 56.9 | 70.3 | |
| LLaVA-1.6Language Model=Vicuna-13B, Prompt Method=MMD-Agent, Parameter Scale=13B2026.03 | 51.8 | 66.7 | 54.6 | 71.4 | |
| LLaVA-1.6Language Model=Vicuna-13B, Prompt=MMD-Agent, Zero-shot=true2026.03 | 51.8 | 66.7 | 54.6 | 71.4 | |
| BLIP2Language Model=FlanT5-XXL, Prompt=MMD-Agent, Zero-shot=true2026.03 | 51.5 | 53.4 | 54 | 53.6 | |
| InstructBLIPLanguage Model=Vicuna-13B, Prompt Method=MMD-Agent, Parameter Scale=13B2026.03 | 51.3 | 53.4 | 54 | 53.1 | |
| InstructBLIPLanguage Model=Vicuna-13B, Prompt=MMD-Agent, Zero-shot=true2026.03 | 51.3 | 53.4 | 54 | 53.1 | |
| LLaVA-1.6Language Model=Vicuna-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 48.1 | 48.2 | 48.5 | 59.5 | |
| LLaVA-1.6Language Model=Vicuna-7B, Prompt=Standard, Zero-shot=true2026.03 | 48.1 | 48.2 | 48.5 | 59.5 | |
| mPLUG-Owl2Language Model=LLaMA2-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 47.2 | 64.9 | 52.3 | 70.6 | |
| mPLUG-Owl2Language Model=LLaMA2-7B, Prompt=Standard, Zero-shot=true2026.03 | 47.2 | 64.9 | 52.3 | 70.6 | |
| Qwen-VLLanguage Model=Qwen-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 43.6 | 50.6 | 44.9 | 60.3 | |
| Qwen-VLLanguage Model=Qwen-7B, Prompt=Standard, Zero-shot=true2026.03 | 43.6 | 50.6 | 44.9 | 60.3 | |
| VILALanguage Model=LLaMA2-7B, Prompt=Standard, Zero-shot=true2026.03 | 41.2 | 35 | 50 | 70 | |
| BLIP2Language Model=FlanT5-XL, Prompt=Standard, Zero-shot=true2026.03 | 41.2 | 35 | 50 | 70 | |
| InstructBLIPLanguage Model=Vicuna-13B, Prompt Method=Standard, Parameter Scale=13B2026.03 | 41.1 | 35 | 49.9 | 69.9 | |
| LLaVA-1.6Language Model=Vicuna-13B, Prompt Method=Standard, Parameter Scale=13B2026.03 | 41.1 | 35 | 50 | 69.7 | |
| VILALanguage Model=LLaMA2-13B, Prompt=Standard, Zero-shot=true2026.03 | 41.1 | 35 | 50 | 70 | |
| InstructBLIPLanguage Model=Vicuna-13B, Prompt=Standard, Zero-shot=true2026.03 | 41.1 | 35 | 49.9 | 69.9 | |
| LLaVA-1.6Language Model=Vicuna-13B, Prompt=Standard, Zero-shot=true2026.03 | 41.1 | 35 | 50 | 69.7 | |
| MiniGPT4Language Model=Vicuna-7B, Prompt=Standard, Zero-shot=true2026.03 | 40.4 | 38.2 | 45.7 | 63.1 | |
| BLIP2Language Model=FlanT5-XXL, Prompt=Standard, Zero-shot=true2026.03 | 31.6 | 63.4 | 53.6 | 35.5 | |
| PandaGPTLanguage Model=Vicuna-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 24.6 | 60.6 | 50.5 | 30.9 | |
| PandaGPTLanguage Model=Vicuna-7B, Prompt=Standard, Zero-shot=true2026.03 | 24.6 | 60.6 | 50.5 | 30.9 | |
| InstructBLIPLanguage Model=Vicuna-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 14.7 | 30.8 | 13.2 | 8.1 | |
| InstructBLIPLanguage Model=Vicuna-7B, Prompt=Standard, Zero-shot=true2026.03 | 14.7 | 30.8 | 13.2 | 8.1 | |
| Otter-ImageLanguage Model=MPT-7B, Prompt=Standard, Zero-shot=true2026.03 | 7.9 | 4.1 | 4.5 | 7.9 |