Forgery Reasoning on TFR (test)
58.8Avg Reasoning Score (Cosine/Rouge-L/BLEU)TextShield-R1
Evaluation Results
| Method | Links | |
|---|---|---|
| TextShield-R1Fine-tuning=Full training set images2026.02 | 58.8 | |
| Qwen2.5-VL-3BFine-tuning=Full training set images2026.02 | 42.9 | |
| Qwen2.5-VL-7BFine-tuning=Full training set images2026.02 | 42.9 | |
| SIDA*Fine-tuning=Full training set images, Base MLLM=Qwen2.5-VL-7B2026.02 | 42.9 | |
| FakeShield*Fine-tuning=Full training set images, Base MLLM=Qwen2.5-VL-7B2026.02 | 42.8 | |
| InternVL3-8BFine-tuning=Full training set images2026.02 | 41.7 | |
| MiniCPM_V_2.6Fine-tuning=Full training set images2026.02 | 41.1 | |
| InternVL3-2BFine-tuning=Full training set images2026.02 | 40.6 | |
| SIDAFine-tuning=Full training set images2026.02 | 35.7 | |
| FakeShieldFine-tuning=Full training set images2026.02 | 35.6 | |
| GPT4oFine-tuning=None2026.02 | 19.4 | |
| InternVL3-8BFine-tuning=None2026.02 | 17.9 | |
| Qwen2.5-VL-3BFine-tuning=None2026.02 | 9.5 | |
| Qwen2.5-VL-7BFine-tuning=None2026.02 | 9.5 | |
| InternVL3-2BFine-tuning=None2026.02 | 8.5 | |
| MiniCPM_V_2.6Fine-tuning=None2026.02 | 3.2 |