Binary manipulation detection on MMFakeBench 10000 samples (test)
74.1F1 ScoreREFORM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| REFORMLanguage Model=Florence2-0.3B, Prompt Method=Ours (Sec.3.2)2026.03 | 74.1 | 74.1 | 74 | 63.7 | |
| REFORMLanguage Model=Florence2-0.3B, Prompt=Ours (Sec.3.2), Zero-shot=true2026.03 | 74.1 | 74.1 | 74 | 63.7 | |
| VILALanguage Model=LLaMA2-13B, Prompt=MMD-Agent, Zero-shot=true2026.03 | 56.6 | 64.3 | 57.2 | 71.2 | |
| LLaVA-1.6Language Model=Vicuna-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 52.5 | 53 | 52.6 | 62.5 | |
| LLaVA-1.6Language Model=Vicuna-7B, Prompt=Standard, Zero-shot=true2026.03 | 52.5 | 53 | 52.6 | 62.5 | |
| BLIP2Language Model=FlanT5-XXL, Prompt=MMD-Agent, Zero-shot=true2026.03 | 51.8 | 54 | 54.7 | 53.5 | |
| LLaVA-1.6Language Model=Vicuna-13B, Prompt Method=MMD-Agent, Parameter Scale=13B2026.03 | 50.2 | 67.3 | 53.9 | 71.3 | |
| LLaVA-1.6Language Model=Vicuna-13B, Prompt=MMD-Agent, Zero-shot=true2026.03 | 50.2 | 67.3 | 53.9 | 71.3 | |
| mPLUG-Owl2Language Model=LLaMA2-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 48.7 | 71.1 | 53.3 | 71.4 | |
| mPLUG-Owl2Language Model=LLaMA2-7B, Prompt=Standard, Zero-shot=true2026.03 | 48.7 | 71.1 | 53.3 | 71.4 | |
| InstructBLIPLanguage Model=Vicuna-13B, Prompt Method=MMD-Agent, Parameter Scale=13B2026.03 | 47.9 | 50.1 | 50.1 | 49.9 | |
| InstructBLIPLanguage Model=Vicuna-13B, Prompt=MMD-Agent, Zero-shot=true2026.03 | 47.9 | 50.1 | 50.1 | 49.9 | |
| Qwen-VLLanguage Model=Qwen-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 44 | 51.6 | 45.2 | 60.5 | |
| Qwen-VLLanguage Model=Qwen-7B, Prompt=Standard, Zero-shot=true2026.03 | 44 | 51.6 | 45.2 | 60.5 | |
| LLaVA-1.6Language Model=Vicuna-13B, Prompt Method=Standard, Parameter Scale=13B2026.03 | 42.3 | 57.3 | 50.1 | 69.5 | |
| LLaVA-1.6Language Model=Vicuna-13B, Prompt=Standard, Zero-shot=true2026.03 | 42.3 | 57.3 | 50.1 | 69.5 | |
| MiniGPT4Language Model=Vicuna-7B, Prompt=Standard, Zero-shot=true2026.03 | 41.7 | 41 | 47.4 | 65.2 | |
| VILALanguage Model=LLaMA2-7B, Prompt=Standard, Zero-shot=true2026.03 | 41.2 | 35 | 50 | 70 | |
| BLIP2Language Model=FlanT5-XL, Prompt=Standard, Zero-shot=true2026.03 | 41.2 | 35 | 50 | 70 | |
| InstructBLIPLanguage Model=Vicuna-13B, Prompt Method=Standard, Parameter Scale=13B2026.03 | 41.1 | 35 | 49.9 | 69.8 | |
| VILALanguage Model=LLaMA2-13B, Prompt=Standard, Zero-shot=true2026.03 | 41.1 | 35 | 50 | 70 | |
| InstructBLIPLanguage Model=Vicuna-13B, Prompt=Standard, Zero-shot=true2026.03 | 41.1 | 35 | 49.9 | 69.8 | |
| BLIP2Language Model=FlanT5-XXL, Prompt=Standard, Zero-shot=true2026.03 | 30.6 | 64.9 | 53.4 | 34.9 | |
| PandaGPTLanguage Model=Vicuna-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 24.1 | 61.7 | 50.4 | 30.6 | |
| PandaGPTLanguage Model=Vicuna-7B, Prompt=Standard, Zero-shot=true2026.03 | 24.1 | 61.7 | 50.4 | 30.6 | |
| InstructBLIPLanguage Model=Vicuna-7B, Prompt Method=Standard, Parameter Scale=7B2026.03 | 16.1 | 40.5 | 14.2 | 8.8 | |
| InstructBLIPLanguage Model=Vicuna-7B, Prompt=Standard, Zero-shot=true2026.03 | 16.1 | 40.5 | 14.2 | 8.8 | |
| Otter-ImageLanguage Model=MPT-7B, Prompt=Standard, Zero-shot=true2026.03 | 8.6 | 32.4 | 5 | 8.6 |