Artifact Explanation on ArtiBench (test)
23.3ROUGEQwen2.5-VL-7B + ArtiAgent
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-VL-7B + ArtiAgentFine-tuning=100K training set generated by ArtiAgent2026.02 | 23.3 | 64.3 | |
| InternVL3.5-8B + ArtiAgentFine-tuning=100K training set generated by ArtiAgent2026.02 | 22.6 | 62.5 | |
| Gemini-2.5-Pro2026.02 | 15.9 | 42 | |
| GPT-52026.02 | 14.5 | 43.4 | |
| LEGIONTraining Split=SynthScars training dataset2026.02 | 14.3 | 33.2 | |
| GPT-4o2026.02 | 14.3 | 43.3 | |
| InternVL3.5-8BFine-tuning=Vanilla2026.02 | 12.6 | 25.6 | |
| Qwen2.5-VL-7BFine-tuning=Vanilla2026.02 | 11.7 | 26.3 |