Visual Question Answering on MMHalSnowball HalluConv 1.0
52AccuracyGPT-4V
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-4VPrompt Type=Formatting Prompt, Model Scale=Closed-Source, Zero-shot=true2024.06 | 52 | 8.49 | 23.3 | 27.69 | |
| GPT-4VPrompt Type=Question Prompt, Model Scale=Closed-Source, Zero-shot=true2024.06 | 42.09 | 9.93 | 14.26 | 43.95 | |
| Qwen-VL-ChatPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 26.2 | 25.6 | 72.48 | 77.83 | |
| MiniGPT-v2Prompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 25.14 | 34.1 | 58.08 | 63.92 | |
| MiniGPT-v2Prompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 21.4 | 40.72 | 66.11 | 72.06 | |
| Qwen-VL-ChatPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 20.03 | 57.91 | 71.7 | 74.97 | |
| ShareGPT4VPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 15.91 | 55.9 | 77.18 | 80.12 | |
| LLaVA-1.5Prompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 14.96 | 56.28 | 78.21 | 81.29 | |
| LLaVA-1.5Prompt Type=Formatting Prompt, Model Scale=13B, Zero-shot=true2024.06 | 14.74 | 57.33 | 78.21 | 81.45 | |
| OtterPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 13.94 | 38.18 | 73.5 | 82.21 | |
| ShareGPT4VPrompt Type=Formatting Prompt, Model Scale=13B, Zero-shot=true2024.06 | 13.43 | 59 | 80.01 | 83.29 | |
| MiniGPT-4Prompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 13.11 | 20.49 | 76.42 | 86.24 | |
| InstructBLIPPrompt Type=Formatting Prompt, Model Scale=13B, Zero-shot=true2024.06 | 12.75 | 40.78 | 76.15 | 85.8 | |
| ShareGPT4VPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 10.54 | 51.27 | 78.01 | 86.27 | |
| LLaVA-1.5Prompt Type=Question Prompt, Model Scale=13B, Zero-shot=true2024.06 | 9.57 | 52.46 | 78.61 | 86.29 | |
| OtterPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 9.43 | 35.47 | 71.61 | 87.42 | |
| mPLUG-Owl2Prompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 7.82 | 52.65 | 86.63 | 89.82 | |
| LLaVA-1.5Prompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 7.68 | 53.53 | 79.96 | 89.03 | |
| IDEFICSPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 7.32 | 33.62 | 85.07 | 91.11 | |
| ShareGPT4VPrompt Type=Question Prompt, Model Scale=13B, Zero-shot=true2024.06 | 6.92 | 57.79 | 83.84 | 90.77 | |
| InstructBLIPPrompt Type=Question Prompt, Model Scale=13B, Zero-shot=true2024.06 | 6.21 | 48.81 | 76.94 | 92.76 | |
| InternLM-XCPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 5.83 | 37.68 | 86.55 | 91.31 | |
| MiniGPT-4Prompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 5.75 | 31.37 | 84.18 | 89.65 | |
| InternLM-XCPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 5.21 | 35.63 | 83.95 | 92.52 | |
| IDEFICSPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 5.05 | 36.17 | 83.37 | 92.83 | |
| mPLUG-Owl2Prompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 4.75 | 50.13 | 84.65 | 93.55 | |
| InstructBLIPPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 4.54 | 55.34 | 90.36 | 93.92 | |
| InstructBLIPPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 4.32 | 56.29 | 85.73 | 94.06 | |
| mPLUG-OwlPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 4.1 | 33.08 | 71.5 | 93.24 | |
| mPLUG-OwlPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 3.64 | 34.16 | 78.62 | 93.35 | |
| CogVLMPrompt Type=Formatting Prompt, Model Scale=7B, Zero-shot=true2024.06 | 2.63 | 72.54 | 93.07 | 96.79 | |
| CogVLMPrompt Type=Question Prompt, Model Scale=7B, Zero-shot=true2024.06 | 2.49 | 70.2 | 92.84 | 96.9 |