Fake Image Detection on HPE-Bench 1.0 (Overall)
95.5AccuracyOurs
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OursFull Method=Layer-Selective MLLM framework2026.01 | 95.5 | 91.62 | |
| AIDECategory=AI-Generated Image Detection Methods, Fine-tuned=true2026.01 | 82.79 | 90.58 | |
| FakeShieldCategory=Deepfake Detection Methods, Fine-tuned=true2026.01 | 73.82 | 85.58 | |
| UnivCategory=AI-Generated Image Detection Methods, Fine-tuned=true2026.01 | 72.5 | 85.5 | |
| LagradCategory=AI-Generated Image Detection Methods, Fine-tuned=true2026.01 | 66.47 | 80.65 | |
| Qwen3-VLCategory=Multimodal Large Language Models, Model Scale=8B2026.01 | 65.59 | 77.43 | |
| LLaVA-NeXTCategory=Multimodal Large Language Models, Model Scale=8B2026.01 | 65.22 | 65.22 | |
| LLama3.2-VisionCategory=Multimodal Large Language Models, Model Scale=11B2026.01 | 65 | 80.23 | |
| HifiNetCategory=Deepfake Detection Methods, Fine-tuned=true2026.01 | 62.87 | 77.02 | |
| CNNSpotCategory=AI-Generated Image Detection Methods, Fine-tuned=true2026.01 | 62.79 | 79.03 | |
| MVSSNetCategory=Deepfake Detection Methods, Fine-tuned=true2026.01 | 62.57 | 74.74 | |
| DeepSeekVL2Category=Multimodal Large Language Models, Model Scale=small2026.01 | 61.4 | 78.67 | |
| MiniCPM-V2.6Category=Multimodal Large Language Models, Model Scale=8B2026.01 | 60.66 | 75.46 | |
| LLaVA-1.6Category=Multimodal Large Language Models, Model Scale=7B2026.01 | 58.6 | 74.21 | |
| InternVL3.5Category=Multimodal Large Language Models, Model Scale=8B2026.01 | 55.29 | 71.12 | |
| mPLUG-Owl3Category=Multimodal Large Language Models, Model Scale=7B2026.01 | 55.15 | 78.33 | |
| Ovis2.5Category=Multimodal Large Language Models, Model Scale=9B2026.01 | 50.66 | 69.18 | |
| Random Choice2026.01 | 50 | 50 | |
| PSCCNetCategory=Deepfake Detection Methods, Fine-tuned=true2026.01 | 46.32 | 76.47 |