Large Multi-modal Model Evaluation on LLaVA-Bench In-the-Wild v1
65.5Conversational ScoreLLaVA-Plus
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaVA-Plustool_usage=All Tools2023.11 | 65.5 | 56.8 | 79.1 | 69.5 | |
| LLaVA-Plustool_usage=Fly2023.11 | 45.2 | 50.4 | 72.6 | 59.1 | |
| LLaVA2023.11 | 42.6 | 51.9 | 68.9 | 57.1 | |
| LLaVAevaluation_protocol=Tools in Test2023.11 | 40.7 | 48.1 | 51.2 | 47.5 | |
| LLaVA-Plustool_usage=Fly, reasoning_thoughts=false2023.11 | 38.8 | 39.8 | 59.8 | 48.7 | |
| GPT4Tools2023.11 | 31.1 | 27.1 | 54.1 | 40.7 |