Pointwise Reasoning Evaluation on DeepfakeJudge Meta-Human
0.5RMSEDeepfakeJudge-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DeepfakeJudge-7BType=Ours2026.02 | 0.5 | 0.25 | 0.96 | 0.95 | |
| DeepfakeJudge-3BType=Ours2026.02 | 0.56 | 0.31 | 0.96 | 0.95 | |
| GPT-4o-MiniType=Closed2026.02 | 0.81 | 0.66 | 0.88 | 0.86 | |
| Qwen-3-VL-235B-ThinkingType=Thinking2026.02 | 0.95 | 0.91 | 0.85 | 0.86 | |
| Qwen-3-VL-30B-ThinkingType=Thinking2026.02 | 1.04 | 1.07 | 0.85 | 0.83 | |
| Gemini-Flash-2.5Type=Closed2026.02 | 1.11 | 1.24 | 0.85 | 0.83 | |
| Qwen-3-VL-8B-ThinkingType=Thinking2026.02 | 1.18 | 1.39 | 0.84 | 0.81 | |
| Qwen-3-VL-235B-InstructType=Open2026.02 | 1.26 | 1.58 | 0.81 | 0.77 | |
| Qwen-3-VL-8B-InstructType=Open2026.02 | 1.28 | 1.63 | 0.79 | 0.76 | |
| Qwen-3-VL-30B-InstructType=Open2026.02 | 1.32 | 1.75 | 0.89 | 0.84 | |
| Qwen-3-VL-2B-InstructType=Open2026.02 | 1.41 | 1.98 | 0.69 | 0.72 | |
| Qwen-3-VL-4B-InstructType=Open2026.02 | 1.53 | 2.34 | 0.6 | 0.61 |