Visual Difference Discovery on VisDiffBench PIS-Hard
61Top-1 AccVisDiff (Caption-based Proposer + Feature-based Ranker)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VisDiff (Caption-based Proposer + Feature-based Ranker)Proposer=Caption (BLIP-2 + GPT-4), Ranker=Feature (CLIP)2023.12 | 61 | 80 | |
| Image-based Proposer (GPT-4V) + Feature-based Ranker (CLIP)Proposer=Image (GPT-4V), Ranker=Feature (CLIP)2023.12 | 57 | 74 | |
| Caption-based Proposer (BLIP-2 + GPT-4) + Image-based Ranker (LLaVA-1.5)Proposer=Caption (BLIP-2 + GPT-4), Ranker=Image (LLaVA-1.5)2023.12 | 38 | 62 | |
| Caption-based Proposer (BLIP-2 + GPT-4) + Caption-based Ranker (Vicuna-1.5)Proposer=Caption (BLIP-2 + GPT-4), Ranker=Caption (Vicuna-1.5)2023.12 | 31 | 61 | |
| Image-based Proposer (LLaVA-1.5) + Feature-based Ranker (CLIP)Proposer=Image (LLaVA-1.5), Ranker=Feature (CLIP)2023.12 | 28 | 43 | |
| Feature-based Proposer (BLIP-2) + Feature-based Ranker (CLIP)Proposer=Feature (BLIP-2), Ranker=Feature (CLIP)2023.12 | 12 | 23 |