Vision-Language Perception and Reasoning on MMAGIBench
65.5Average AccuracyOtter
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| OtterLanguage Decoder=LLAMA-7B2023.06 | 65.5 | 68.9 | 47.3 | 66.3 | 61.8 | 83.3 | |
| LLaVALanguage Decoder=Vicuna-7B2023.06 | 62.7 | 44.4 | 54.2 | 71.9 | 76.5 | 66.7 | |
| OpenFlamingoLanguage Decoder=LLAMA-7B2023.06 | 51.1 | 34.4 | 40 | 61.3 | 52.9 | 66.7 | |
| MiniGPT-4Language Decoder=Vicuna-7B2023.06 | 51 | 63.3 | 47.8 | 50.6 | 26.5 | 66.7 | |
| InstructBLIPLanguage Decoder=Vicuna-7B2023.06 | 50.4 | 67.8 | 52.2 | 43.8 | 38.2 | 50 |