Large Multi-modal Model Evaluation on LLaVA-Bench COCO v1
0.82Conv ScoreLLaVA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaVA2023.11 | 0.82 | 0.691 | 0.926 | 0.812 | |
| LLaVA-Plustool_usage=All Tools2023.11 | 0.816 | 0.745 | 0.957 | 0.839 | |
| LLaVA-Plustool_usage=Fly, reasoning_thoughts=false2023.11 | 0.766 | 0.704 | 0.907 | 0.794 | |
| LLaVA-Plustool_usage=Fly2023.11 | 0.762 | 0.722 | 0.923 | 0.804 | |
| GPT4Tools2023.11 | 0.753 | 0.538 | 0.869 | 0.721 | |
| LLaVAevaluation_protocol=Tools in Test2023.11 | 0.562 | 0.679 | 0.533 | 0.591 |