Visual Question Answering on MMHalSnowball CleanConv 1.0
77.94AccuracyQwen-VL-Chat
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-VL-ChatPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 77.94 | |
| CogVLMPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 75.17 | |
| CogVLMPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 72.69 | |
| ShareGPT4VPrompt Type=Formatting Prompt, Model Scale=13B, Zero-shot=true2024.06 | 72.43 | |
| LLaVA-1.5Prompt Type=Formatting Prompt, Model Scale=13B, Zero-shot=true2024.06 | 72.07 | |
| ShareGPT4VPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 71.81 | |
| LLaVA-1.5Prompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 71.24 | |
| ShareGPT4VPrompt Type=Question Prompt, Model Scale=13B, Zero-shot=true2024.06 | 64.71 | |
| MiniGPT-v2Prompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 62.12 | |
| LLaVA-1.5Prompt Type=Question Prompt, Model Scale=13B, Zero-shot=true2024.06 | 62.03 | |
| ShareGPT4VPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 61.81 | |
| LLaVA-1.5Prompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 61.21 | |
| InstructBLIPPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 60.61 | |
| GPT-4VPrompt Type=Formatting Prompt, Model Scale=Closed-Source, Zero-shot=true2024.06 | 60.49 | |
| mPLUG-Owl2Prompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 60.47 | |
| InstructBLIPPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 59.88 | |
| MiniGPT-v2Prompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 59.24 | |
| InstructBLIPPrompt Type=Question Prompt, Model Scale=13B, Zero-shot=true2024.06 | 55.02 | |
| mPLUG-Owl2Prompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 54.88 | |
| InstructBLIPPrompt Type=Formatting Prompt, Model Scale=13B, Zero-shot=true2024.06 | 53.53 | |
| OtterPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 52.12 | |
| GPT-4VPrompt Type=Question Prompt, Model Scale=Closed-Source, Zero-shot=true2024.06 | 52.02 | |
| Qwen-VL-ChatPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 51.8 | |
| OtterPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 44.9 | |
| InternLM-XCPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 43.51 | |
| IDEFICSPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 41.22 | |
| IDEFICSPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 40.94 | |
| InternLM-XCPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 40.84 | |
| mPLUG-OwlPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 37.8 | |
| mPLUG-OwlPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 37.18 | |
| MiniGPT-4Prompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 37.12 | |
| MiniGPT-4Prompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 33.6 |