Multimodal Instruction Following on ViSiT-Bench Sept. 27th, 2023 (Leaderboard)
1,382ELOHuman Reference
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Human Reference2023.11 | 1,382 | 5,880 | — | — | |
| LLaVA-PlusModel Size=13B2023.11 | 1,203 | 678 | 35.07 | 134 | |
| LLaVAModel Size=13B2023.11 | 1,095 | 5,420 | 18.53 | 475 | |
| mPLUG-Owl2023.11 | 1,087 | 5,440 | 15.83 | 480 | |
| LlamaAdapter-v22023.11 | 1,066 | 5,469 | 14.14 | 488 | |
| LynxModel Size=8B2023.11 | 1,037 | 787 | 11.43 | 140 | |
| IdeficsModel Size=9B2023.11 | 1,020 | 794 | 9.72 | 144 | |
| InstructBLIP2023.11 | 1,000 | 5,469 | 14.12 | 503 | |
| Otter2023.11 | 962 | 5,443 | 7.01 | 499 | |
| Visual GPT2023.11 | 941 | 5,437 | 1.57 | 510 | |
| MiniGPT-42023.11 | 926 | 5,448 | 3.36 | 506 | |
| Octopus V22023.11 | 925 | 790 | 8.9 | 146 | |
| OpenFlamingo V12023.11 | 851 | 5,479 | 2.95 | 509 | |
| PandaGPTModel Size=13B2023.11 | 775 | 5,465 | 2.7 | 519 | |
| MultimodalGPT2023.11 | 731 | 5,471 | 0.19 | 527 |