Visual Question Answering on OK-VQA
80.7VQA ScoreLAMP
Evaluation Results
| Method | Links | |
|---|---|---|
| LAMPModel=MiniGPT42026.01 | 80.7 | |
| LAMPModel=LLaVA2026.01 | 74.5 | |
| LAMPModel=Mantis-CLIP2026.01 | 73.4 | |
| OMGM2025.05 | 66.57 | |
| PreFLMR2025.05 | 61.88 | |
| PICaModel size=175B, shots=162021.10 | 43.3 | |
| X-transferModel=LLaVA2026.01 | 31.5 | |
| X-transferModel=MiniGPT2026.01 | 30.1 | |
| X-transferModel=Mantis-CLIP2026.01 | 29.5 | |
| FEWVLM_largeModel size=740M, shots=162021.10 | 23.1 | |
| METALMk (number of in-context examples)=4, learning_mode=in-context learning, prediction_manner=generative2022.06 | 16 | |
| FEWVLM_baseModel size=224M, shots=162021.10 | 15 | |
| FEWVLM_baseModel size=224M, shots=42021.10 | 14.5 | |
| METALMk (number of in-context examples)=1, learning_mode=in-context learning, prediction_manner=generative2022.06 | 13.2 | |
| VL-T5no-vqaModel size=224M, shots=162021.10 | 12.7 | |
| FrozenModel size=7B, shots=42021.10 | 12.6 | |
| Frozenk (number of in-context examples)=4, learning_mode=in-context learning, prediction_manner=generative2022.06 | 12.6 | |
| Frozenk (number of in-context examples)=1, learning_mode=in-context learning, prediction_manner=generative2022.06 | 9.7 |