Deepfake Detection on DeepfakeJudge-Detect (test)
96.6Accuracy (Real)Gemini-2.5-Flash
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Gemini-2.5-FlashType=Closed2026.02 | 96.6 | 73.7 | 34.5 | 50 | 65.5 | 61.9 | |
| ChatGPT-4o-miniType=Closed2026.02 | 95.8 | 70.2 | 22.7 | 35.8 | 59.3 | 53 | |
| Qwen-3-VL-30BType=Open2026.02 | 94.6 | 74.5 | 41 | 56 | 67.7 | 65.3 | |
| Qwen-3-VL-235BType=Open2026.02 | 93.5 | 78.6 | 55.4 | 68.4 | 74.5 | 73.5 | |
| Qwen-3-VL-235B-ThinkingType=Reasoning2026.02 | 75 | 76.6 | 90.3 | 79.8 | 63.7 | 73.4 | |
| SIDA-13B-DescriptionType=DF2026.02 | 67.6 | 57 | 27.9 | 34.5 | 48.1 | 45.8 | |
| Qwen-3-VL-30B-ThinkingType=Reasoning2026.02 | 67.4 | 66 | 87.6 | 72.9 | 47.2 | 59.1 | |
| Qwen-3-VL-8B-ThinkingType=Reasoning2026.02 | 67.1 | 67.1 | 78.7 | 69.9 | 55.5 | 64.3 | |
| Microsoft-Phi-4-InstructType=Open2026.02 | 61 | 60.8 | 54.4 | 58.2 | 67.5 | 63.4 | |
| Google-Gemma-12BType=Open2026.02 | 57.7 | 57.4 | 49.4 | 54 | 66.1 | 60.8 | |
| InternVL3.5-GPT-OSS-20B-A4BType=Open2026.02 | 55.6 | 67.6 | 55.6 | 29.2 | 55.6 | 48.4 | |
| Qwen-3-VL-8B-InstructType=Open2026.02 | 50.4 | 23.7 | 50.4 | 63.2 | 50.4 | 43.5 | |
| Qwen2.5-VL-Gen-Buster++Type=DF2026.02 | 49.9 | 40 | 49.9 | 66.5 | 49.9 | 33.5 | |
| Qwen3-VL-2B-InstructType=Open2026.02 | 49.8 | 44.7 | 49.8 | 54 | 49.8 | 49.3 | |
| InternVL3.5-1B-HFType=Open2026.02 | 47.8 | 63 | 47.8 | 11.4 | 47.8 | 37.2 |