Pairwise Comparison on DeepfakeJudge Meta-Human
99.4Pairwise AccuracyQwen-3-VL-235B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-3-VL-235B-InstructType=Open2026.02 | 99.4 | |
| DeepfakeJudge-7BType=Ours2026.02 | 98.9 | |
| Qwen-3-VL-30B-ThinkingType=Thinking2026.02 | 97.7 | |
| DeepfakeJudge-3BType=Ours2026.02 | 96.6 | |
| Qwen-3-VL-30B-InstructType=Open2026.02 | 96.3 | |
| Qwen-3-VL-235B-ThinkingType=Thinking2026.02 | 95.5 | |
| Gemini-Flash-2.5Type=Closed2026.02 | 94.2 | |
| Qwen-3-VL-8B-ThinkingType=Thinking2026.02 | 93.2 | |
| GPT-4o-MiniType=Closed2026.02 | 89.8 | |
| Qwen-3-VL-8B-InstructType=Open2026.02 | 88.6 | |
| Qwen-3-VL-4B-InstructType=Open2026.02 | 72.7 | |
| Qwen-3-VL-2B-InstructType=Open2026.02 | 65.1 |