Pointwise Reasoning Evaluation on DeepfakeJudge Meta
0.61RMSEDeepfakeJudge-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DeepfakeJudge-7BType=Ours2026.02 | 0.61 | 0.37 | 0.94 | 0.93 | |
| DeepfakeJudge-3BType=Ours2026.02 | 0.69 | 0.48 | 0.92 | 0.92 | |
| GPT-4o-MiniType=Closed2026.02 | 0.78 | 0.6 | 0.87 | 0.87 | |
| Qwen-3-VL-235B-ThinkingType=Thinking2026.02 | 1.01 | 1.02 | 0.84 | 0.84 | |
| Qwen-3-VL-30B-ThinkingType=Thinking2026.02 | 1.02 | 1.05 | 0.83 | 0.83 | |
| Gemini-Flash-2.5Type=Closed2026.02 | 1.09 | 1.2 | 0.82 | 0.83 | |
| Qwen-3-VL-8B-ThinkingType=Thinking2026.02 | 1.09 | 1.19 | 0.83 | 0.82 | |
| Qwen-3-VL-235B-InstructType=Open2026.02 | 1.1 | 1.21 | 0.83 | 0.82 | |
| Qwen-3-VL-8B-InstructType=Open2026.02 | 1.19 | 1.41 | 0.76 | 0.75 | |
| Qwen-3-VL-30B-InstructType=Open2026.02 | 1.28 | 1.65 | 0.81 | 0.81 | |
| Qwen-3-VL-4B-InstructType=Open2026.02 | 1.59 | 2.53 | 0.57 | 0.58 | |
| Qwen-3-VL-2B-InstructType=Open2026.02 | 1.6 | 2.55 | 0.6 | 0.61 |