Visual Difference Discovery on VisDiffBench ImageNet-R/*
86Acc@1Image-based Proposer (GPT-4V) + Feature-based Ranker (CLIP)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Image-based Proposer (GPT-4V) + Feature-based Ranker (CLIP)Proposer=Image (GPT-4V), Ranker=Feature (CLIP)2023.12 | 86 | 92 | |
| Caption-based Proposer (BLIP-2 + GPT-4) + Image-based Ranker (LLaVA-1.5)Proposer=Caption (BLIP-2 + GPT-4), Ranker=Image (LLaVA-1.5)2023.12 | 78 | 88 | |
| VisDiff (Caption-based Proposer + Feature-based Ranker)Proposer=Caption (BLIP-2 + GPT-4), Ranker=Feature (CLIP)2023.12 | 78 | 96 | |
| Feature-based Proposer (BLIP-2) + Feature-based Ranker (CLIP)Proposer=Feature (BLIP-2), Ranker=Feature (CLIP)2023.12 | 68 | 85 | |
| Caption-based Proposer (BLIP-2 + GPT-4) + Caption-based Ranker (Vicuna-1.5)Proposer=Caption (BLIP-2 + GPT-4), Ranker=Caption (Vicuna-1.5)2023.12 | 42 | 70 | |
| Image-based Proposer (LLaVA-1.5) + Feature-based Ranker (CLIP)Proposer=Image (LLaVA-1.5), Ranker=Feature (CLIP)2023.12 | 27 | 39 |