Multimodal Instruction Following on LLaVA-Bench In-the-Wild 1.0 (test)
58.8ConversationLLaVA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaVAevaluation_protocol=GPT-4 queried three times for consistent evaluation on single set2023.04 | 58.8 | 49.2 | 81.4 | 66.7 | |
| LLaVAevaluation_protocol=mean of three inference runs2023.04 | 57.3 | 52.5 | 81.7 | 67.3 | |
| BLIP-2evaluation_protocol=mean of three inference runs2023.04 | 54.6 | 29.1 | 32.9 | 38.1 | |
| OpenFlamingoevaluation_protocol=mean of three inference runs2023.04 | 19.3 | 19 | 19.1 | 19.1 |