Visual Question Answering on OKVQA (val)
66.1VQA ScorePaLM-E
Evaluation Results
| Method | Links | |
|---|---|---|
| PaLM-EParameters=562B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 66.1 | |
| PaLI-XParameters=55B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 66.1 | |
| PaLM-ENumber of parameters=562B2023.05 | 66.1 | |
| PaLI-XNumber of parameters=55B2023.05 | 66.1 | |
| PaLM-EParameters=562B2023.10 | 66.1 | |
| PaLI-XParameters=55B2023.10 | 66.1 | |
| PalmE-562Btrained_on_QA_datasets=true2024.03 | 66.1 | |
| Emu2-ChatLLM=LLaMA-33B, Trained during SFT stage=true2023.11 | 64.8 | |
| CogVLM-ChatLLM=Vicuna-7B, Trained during SFT stage=true2023.11 | 64.8 | |
| multi-modal exploration-exploitation reinforcement learning frameworkBase Model=Qwen-VL, Selection Strategy=Reinforcement Learning2025.06 | 64.7 | |
| PaLIParameters=17B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 64.5 | |
| PaLINumber of parameters=17B2023.05 | 64.5 | |
| PaLI-17BParameters=17B2023.10 | 64.5 | |
| PaLI-17Btrained_on_QA_datasets=true2024.03 | 64.5 | |
| SPHINX-2kLLM=LLaMA2 13B, Trained during SFT stage=true2023.11 | 62.6 | |
| randomBase Model=Qwen-VL, Selection Strategy=random2025.06 | 61.1 | |
| similarityBase Model=Qwen-VL, Selection Strategy=Similarity2025.06 | 60.6 | |
| De-DiffusionVisual Encoder=ViT-L, LLM=PaLM 2-L, Trainable Parameters=135M, Shots=322023.11 | 60.6 | |
| FuyuLLM=Fuyu-8B, Trained during SFT stage=true2023.11 | 60.6 | |
| BM25Base Model=Qwen-VL, Selection Strategy=BM252025.06 | 60.3 | |
| PaLI-3Parameters=5B2023.10 | 60.1 | |
| De-DiffusionVisual Encoder=ViT-L, LLM=PaLM 2-L, Trainable Parameters=135M, Shots=42023.11 | 58.2 | |
| FlamingoParameters=80B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 57.8 | |
| FlamingoNumber of parameters=80B2023.05 | 57.8 | |
| FlamingoParameters=80B, Shots=322023.10 | 57.8 | |
| IDEFICS-80BLLM=Llama65B, Trainable Parameters=14B, Shots=322023.11 | 57.8 | |
| Flamingo-80BLLM=Chinchilla70B, Trainable Parameters=10B, Shots=322023.11 | 57.8 | |
| mPLUG-Owl2LLM=LLaMA2-7B, Trained during SFT stage=true2023.11 | 57.7 | |
| zero-shotBase Model=Qwen-VL, Selection Strategy=zero-shot2025.06 | 57.5 | |
| Flamingo-80BLLM=Chinchilla70B, Trainable Parameters=10B, Shots=42023.11 | 57.4 | |
| De-DiffusionVisual Encoder=ViT-L, LLM=PaLM 2-L, Trainable Parameters=135M, Shots=02023.11 | 57 | |
| Qwen-VL-ChatLLM=Qwen-7B, Trained during SFT stage=true2023.11 | 56.6 | |
| Unified-IO2LLM=UIO-2XXL, Trained during SFT stage=true2023.11 | 55.5 | |
| PalmE-12Btrained_on_QA_datasets=true2024.03 | 55.5 | |
| Idefics2-baseSize=8B, Architecture=fully autoregressive, # tokens per image=64, Few-shot=8 random in-context examples, Setting=open-ended2024.05 | 54.6 | |
| De-DiffusionVisual Encoder=ViT-L, LLM=PaLM 2-S, Trainable Parameters=135M, Shots=42023.11 | 53.5 | |
| De-DiffusionVisual Encoder=ViT-L, LLM=PaLM 2-S, Trainable Parameters=135M, Shots=322023.11 | 53.3 | |
| IDEFICS-80BLLM=Llama65B, Trainable Parameters=14B, Shots=42023.11 | 52.4 | |
| PaLI-3Btrained_on_QA_datasets=true2024.03 | 52.4 | |
| FlexCap-LLMzero-shot=true2024.03 | 52.1 | |
| ViperGPTzero-shot=true2024.03 | 51.9 | |
| De-DiffusionVisual Encoder=ViT-L, LLM=PaLM 2-S, Trainable Parameters=135M, Shots=02023.11 | 51.4 | |
| MM1Size=7B, Architecture=fully autoregressive, # tokens per image=144, Few-shot=8 random in-context examples, Setting=open-ended2024.05 | 51.4 | |
| Flamingo-9BLLM=Chinchilla7B, Trainable Parameters=2B, Shots=322023.11 | 51 | |
| Flamingo-80BLLM=Chinchilla70B, Trainable Parameters=10B, Shots=02023.11 | 50.6 | |
| Flamingozero-shot=true2024.03 | 50.6 | |
| FlamingoSize=9B, Architecture=cross-attention, Few-shot=8 random in-context examples, Setting=open-ended2024.05 | 50 | |
| Flamingo-9BLLM=Chinchilla7B, Trainable Parameters=2B, Shots=42023.11 | 49.3 | |
| PICa-FullLLM=GPT-3, Trainable Parameters=0, Shots=162023.11 | 48 | |
| Idefics1Size=9B, Architecture=cross-attention, Few-shot=8 random in-context examples, Setting=open-ended2024.05 | 47.7 | |
| BLIP-2Visual Encoder=ViT-g, LLM=FlanT5XXL, Trainable Parameters=108M, Shots=0, In-domain COCO training=true2023.11 | 45.9 | |
| BLIPv2zero-shot=true2024.03 | 45.9 | |
| IDEFICS-80BLLM=Llama65B, Trainable Parameters=14B, Shots=02023.11 | 45.2 | |
| Flamingo-9BLLM=Chinchilla7B, Trainable Parameters=2B, Shots=02023.11 | 44.7 | |
| multi-modal exploration-exploitation reinforcement learning frameworkBase Model=LLaVA, Selection Strategy=Reinforcement Learning2025.06 | 44.3 | |
| DreamLLMLLM=Vicuna-7B, Trained during SFT stage=false2023.11 | 44.3 | |
| LENSLLM=FlanT5XXL, Trainable Parameters=0, Shots=02023.11 | 43.3 | |
| AnyMALVisual Encoder=VIT-G, LLM=Llama270B, Trainable Parameters=0, Shots=0, In-domain COCO training=true2023.11 | 42.6 | |
| OpenFlamingo-9BLLM=MPT7B, Shots=322023.11 | 42.4 | |
| OpenFlamingoSize=9B, Architecture=cross-attention, Few-shot=8 random in-context examples, Setting=open-ended2024.05 | 41.1 | |
| OpenFlamingo-9BLLM=MPT7B, Shots=42023.11 | 40.1 | |
| MAVEXuses_task_training_data=true2021.06 | 39.4 | |
| OpenFlamingoLLM=MPT-7B, Trained during SFT stage=false2023.11 | 38.3 | |
| OpenFlamingo-9BLLM=MPT7B, Shots=02023.11 | 37.8 | |
| IDEFICS-InstructLLM=LLaMA-65B, Trained during SFT stage=false2023.11 | 36.9 | |
| VL-HyperPELTNumber of Samples=20002022.03 | 35.86 | |
| VL-HyperPELTNumber of Samples=10002022.03 | 35.72 | |
| VL-HyperPELTNumber of Samples=5002022.03 | 35.56 | |
| VL-HyperPELTNumber of Samples=1002022.03 | 34.99 | |
| VL-AdapterNumber of Samples=20002022.03 | 34.87 | |
| VL-HyperPELTNumber of Samples=322022.03 | 34.86 | |
| VL-HyperPELTNumber of Samples=162022.03 | 34.72 | |
| CLIP-T5Number of Samples=20002022.03 | 34.62 | |
| CLIP-T5Number of Samples=10002022.03 | 34.59 | |
| VL-AdapterNumber of Samples=10002022.03 | 34.57 | |
| CLIP-T5Number of Samples=5002022.03 | 34.43 | |
| CLIP-T5Number of Samples=1002022.03 | 34.27 | |
| CLIP-T5Number of Samples=322022.03 | 33.87 | |
| CLIP-T5Number of Samples=162022.03 | 33.68 | |
| VL-AdapterNumber of Samples=5002022.03 | 33.35 | |
| VL-HyperPELTNumber of Samples=42022.03 | 33.25 | |
| VL-AdapterNumber of Samples=1002022.03 | 33.03 | |
| CLIP-T5Number of Samples=42022.03 | 32.65 | |
| VL-AdapterNumber of Samples=322022.03 | 32.07 | |
| VL-AdapterNumber of Samples=162022.03 | 31.86 | |
| VL-AdapterNumber of Samples=42022.03 | 31.83 | |
| similarityBase Model=LLaVA, Selection Strategy=Similarity2025.06 | 22.7 | |
| Frozen VQAn-shot=0, uses_task_training_data=false2021.06 | 19.6 | |
| Frozenn-shot=4, uses_task_training_data=false2021.06 | 12.6 | |
| Frozen VQA-blindn-shot=0, uses_task_training_data=false2021.06 | 12.5 | |
| Frozenn-shot=1, uses_task_training_data=false2021.06 | 9.7 | |
| Frozen train-blindn-shot=1, uses_task_training_data=false2021.06 | 7.2 | |
| Frozen 400mLMn-shot=4, uses_task_training_data=false2021.06 | 6.6 | |
| Frozenn-shot=0, uses_task_training_data=false2021.06 | 5.9 | |
| Frozen 400mLMn-shot=1, uses_task_training_data=false2021.06 | 5.9 | |
| Frozen finetunedn-shot=4, uses_task_training_data=false2021.06 | 4.6 | |
| Frozen finetunedn-shot=0, uses_task_training_data=false2021.06 | 4.2 | |
| Frozen finetunedn-shot=1, uses_task_training_data=false2021.06 | 4.1 | |
| Frozen 400mLMn-shot=0, uses_task_training_data=false2021.06 | 4 | |
| Frozen train-blindn-shot=0, uses_task_training_data=false2021.06 | 3.3 |