Visual Question Answering on AnimalKB 1.0 (test)
98.68Visual Accgpt-4o (0806)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| gpt-4o (0806)Knowledge Context=with KB2025.12 | 98.68 | 97.79 | 98.24 | |
| gpt-4o (1120)Knowledge Context=with KB2025.12 | 98.26 | 97.87 | 98.06 | |
| gemini-2.0-flashKnowledge Context=with KB2025.12 | 98.16 | 97.65 | 97.9 | |
| claude-3.5-sonnetKnowledge Context=with KB2025.12 | 97.43 | 96.72 | 97.07 | |
| gemini-1.5-proKnowledge Context=with KB2025.12 | 97.01 | 96.54 | 96.78 | |
| gpt-4o-miniKnowledge Context=with KB2025.12 | 94.04 | 95.07 | 94.56 | |
| claude-3-haikuKnowledge Context=with KB2025.12 | 92.06 | 92.21 | 92.13 | |
| InternVL2.5-26BModel Size=25.5 B, Knowledge Context=with KB2025.12 | 87.99 | 88.04 | 88.01 | |
| Qwen2-VL-7BModel Size=8.3 B, Knowledge Context=with KB2025.12 | 87.67 | 86.91 | 87.29 | |
| InternVL2.5-8BModel Size=8.1 B, Knowledge Context=with KB2025.12 | 87.23 | 86.47 | 86.85 | |
| Qwen2.5-VL-7BModel Size=8.3 B, Knowledge Context=with KB2025.12 | 87.06 | 85.42 | 86.24 | |
| DeepSeek-VL2-SmallModel Size=16.1 B, Knowledge Context=with KB2025.12 | 85.39 | 84.73 | 85.06 | |
| gemini-2.0-flashKnowledge Context=without KB2025.12 | 84.95 | 81.96 | 83.46 | |
| DeepSeek-VL2Model Size=27.5 B, Knowledge Context=with KB2025.12 | 84.66 | 85.1 | 84.88 | |
| LLaVA-Next-Llama3-8BModel Size=8.4 B, Knowledge Context=with KB2025.12 | 83.41 | 84.19 | 83.8 | |
| LLaVA-OV-Qwen2-7BModel Size=8.0 B, Knowledge Context=with KB2025.12 | 83.28 | 81.81 | 82.55 | |
| claude-3.5-sonnetKnowledge Context=without KB2025.12 | 81.05 | 77.82 | 79.44 | |
| LLaVA-v1.6-Vicuna-7BModel Size=7.1 B, Knowledge Context=with KB2025.12 | 80.71 | 80.86 | 80.78 | |
| gpt-4o (0806)Knowledge Context=without KB2025.12 | 79.36 | 77.13 | 78.25 | |
| gpt-4o (1120)Knowledge Context=without KB2025.12 | 78.95 | 76.81 | 77.88 | |
| gemini-1.5-proKnowledge Context=without KB2025.12 | 78.68 | 78.68 | 79.71 | |
| InternVL2.5-26BModel Size=25.5 B, Knowledge Context=without KB2025.12 | 76.52 | 72.55 | 74.53 | |
| Qwen2-VL-7BModel Size=8.3 B, Knowledge Context=without KB2025.12 | 74.56 | 72.84 | 73.7 | |
| Qwen2.5-VL-7BModel Size=8.3 B, Knowledge Context=without KB2025.12 | 74.44 | 71.25 | 72.84 | |
| DeepSeek-VL2Model Size=27.5 B, Knowledge Context=without KB2025.12 | 73.75 | 75.69 | 74.72 | |
| LLaVA-OV-Qwen2-7BModel Size=8.0 B, Knowledge Context=without KB2025.12 | 71.94 | 65.27 | 68.6 | |
| InternVL2.5-8BModel Size=8.1 B, Knowledge Context=without KB2025.12 | 71.27 | 69.09 | 70.18 | |
| gpt-4o-miniKnowledge Context=without KB2025.12 | 69.78 | 66.86 | 68.32 | |
| DeepSeek-VL2-SmallModel Size=16.1 B, Knowledge Context=without KB2025.12 | 66.13 | 68.7 | 67.41 | |
| LLaVA-v1.6-Vicuna-7BModel Size=7.1 B, Knowledge Context=without KB2025.12 | 65.76 | 60.91 | 63.33 | |
| LLaVA-Next-Llama3-8BModel Size=8.4 B, Knowledge Context=without KB2025.12 | 65.47 | 63.75 | 64.61 | |
| claude-3-haikuKnowledge Context=without KB2025.12 | 63.82 | 59.14 | 61.48 |