Region-level instance understanding on ViP-Bench Visual prompts from human
57.7RecLLaVA-NeXT-INST-IT-Qwen2-7B
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| LLaVA-NeXT-INST-IT-Qwen2-7BBase Model=Qwen2-7B, Visual prompt source=Human, Evaluation protocol=Zero-shot2024.12 | 57.7 | 22.5 | 53.2 | 19.4 | 53.6 | 45 | 49 | |
| GPT-4V-turbo-detail:highDetail Level=high, Visual prompt source=Human, Evaluation protocol=Zero-shot2024.12 | 56.9 | 69.7 | 63.7 | 80.6 | 61.1 | 45.6 | 59.9 | |
| ViP-LLaVA-7BVisual prompt source=Human, Evaluation protocol=Zero-shot2024.12 | 55.3 | 17.6 | 45.9 | 8.1 | 44.6 | 33.1 | 46.8 | |
| LLaVA-NeXT-INST-IT-Vicuna-7BBase Model=Vicuna-7B, Visual prompt source=Human, Evaluation protocol=Zero-shot2024.12 | 55 | 21.3 | 52.5 | 16.1 | 57.5 | 40.6 | 48.2 | |
| GPT-4V-turbo-detail:lowDetail Level=low, Visual prompt source=Human, Evaluation protocol=Zero-shot2024.12 | 51.7 | 50.3 | 59.3 | 60.3 | 55 | 43.8 | 51.4 | |
| LLAVA-1.5-7BVisual prompt source=Human, Evaluation protocol=Zero-shot2024.12 | 49.1 | 13 | 42.9 | 9.7 | 50 | 27.5 | 40.2 | |
| Qwen-VL-ChatVisual prompt source=Human, Evaluation protocol=Zero-shot2024.12 | 48.7 | 22.1 | 41.2 | 6.5 | 48.2 | 25 | 41.7 | |
| InstructBLIP-7BVisual prompt source=Human, Evaluation protocol=Zero-shot2024.12 | 38.9 | 17 | 35.4 | 9.7 | 29.3 | 17.5 | 33.3 |