General VQA on POPE
89AccuracyMolmo
Evaluation Results
| Method | Links | |
|---|---|---|
| MolmoOpenness=fully open, Model Size=7B, Training Strategy=Distilled (-D)2025.10 | 89 | |
| InternVL3.5Openness=semi-open, Model Size=8B2025.10 | 88.7 | |
| LLaVA OneVisionOpenness=fully open, Model Size=7B2025.10 | 88.4 | |
| TreeVGR-7BParameters=7B2025.07 | 87.2 | |
| Qwen2.5-VL-7BParameters=7B2025.07 | 86.7 | |
| Qwen2.5-VLOpenness=semi-open, Model Size=7B2025.10 | 86.4 | |
| Keye-VLOpenness=semi-open, Model Size=8B2025.10 | 86 | |
| Qwen2.5-VL-72BParameters=72B2025.07 | 84.9 | |
| Bee-8BOpenness=fully open, Model Size=8B, Training Strategy=RL2025.10 | 84.8 | |
| Bee-8BOpenness=fully open, Model Size=8B, Training Strategy=SFT2025.10 | 84 | |
| CoM-PTBackbone=ViT-L/16, PT Method=CoM-PT, Fine-tuning Protocol=LoRA-based LLaVA-1.5-7B [41]2026.04 | 75.15 | |
| BaselineBackbone=ViT-L/16, PT Method=Baseline, Fine-tuning Protocol=LoRA-based LLaVA-1.5-7B [41]2026.04 | 74.5 | |
| CoM-PTBackbone=ViT-B/16, PT Method=CoM-PT, Fine-tuning Protocol=LoRA-based LLaVA-1.5-7B [41]2026.04 | 74.17 | |
| BaselineBackbone=ViT-B/16, PT Method=Baseline, Fine-tuning Protocol=LoRA-based LLaVA-1.5-7B [41]2026.04 | 73.62 |