Pairwise Comparison on DeepfakeJudge Meta
96.2Pairwise AccuracyDeepfakeJudge-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepfakeJudge-7BType=Ours2026.02 | 96.2 | |
| DeepfakeJudge-3BType=Ours2026.02 | 94.4 | |
| Qwen-3-VL-235B-InstructType=Open2026.02 | 93.2 | |
| Qwen-3-VL-30B-ThinkingType=Thinking2026.02 | 92.5 | |
| Gemini-Flash-2.5Type=Closed2026.02 | 91.7 | |
| Qwen-3-VL-30B-InstructType=Open2026.02 | 91.3 | |
| Qwen-3-VL-235B-ThinkingType=Thinking2026.02 | 90.8 | |
| GPT-4o-MiniType=Closed2026.02 | 90.3 | |
| Qwen-3-VL-8B-ThinkingType=Thinking2026.02 | 89.2 | |
| Qwen-3-VL-8B-InstructType=Open2026.02 | 86 | |
| Qwen-3-VL-4B-InstructType=Open2026.02 | 75.8 | |
| Qwen-3-VL-2B-InstructType=Open2026.02 | 74.8 |