Multimodal In-context Learning on Multimodal Benchmarks Average
67.2AccuracyAIMv2
Evaluation Results
| Method | Links | |
|---|---|---|
| AIMv2architecture=ViT-L/14, shots=8-shot2024.11 | 67.2 | |
| DFN-CLIParchitecture=ViT-H/14, shots=8-shot2024.11 | 66.4 | |
| OAI CLIParchitecture=ViT-L/14, shots=8-shot2024.11 | 66.1 | |
| AIMv2architecture=ViT-L/14, shots=4-shot2024.11 | 63.8 | |
| DFN-CLIParchitecture=ViT-H/14, shots=4-shot2024.11 | 62.5 | |
| OAI CLIParchitecture=ViT-L/14, shots=4-shot2024.11 | 62.2 | |
| DFN-CLIParchitecture=ViT-H/14, shots=0-shot2024.11 | 40.9 | |
| AIMv2architecture=ViT-L/14, shots=0-shot2024.11 | 39.6 | |
| OAI CLIParchitecture=ViT-L/14, shots=0-shot2024.11 | 39.3 |