Object Classification on COCO 2017 (val)
82.9AccuracySpatialRGPT-VILA-1.5-8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SpatialRGPT-VILA-1.5-8BEvaluation Protocol=Ground-truth boxes, Backbone=VILA-1.5, Parameters=8B2024.06 | 82.9 | 72.9 | |
| SpatialRGPT-VILA-1.5-3BEvaluation Protocol=Ground-truth boxes, Backbone=VILA-1.5, Parameters=3B2024.06 | 82.5 | 72.5 | |
| RegionGPTPT=923K, IT=953K, Vision Backbone=ViT-L, LLM Backbone=Vicuna-7B2024.03 | 80.61 | 70 | |
| RegionGPT-7BEvaluation Protocol=Ground-truth boxes, Parameters=7B2024.06 | 80.6 | 70 | |
| SpatialRGPT-7BEvaluation Protocol=Ground-truth boxes, Parameters=7B2024.06 | 79.9 | 69.7 | |
| GSCG ClassifierContext Configuration=Full Context2025.12 | 73.43 | — | |
| PVITPT=13.7M, IT=243K, Vision Backbone=ViT-L + R50x4, LLM Backbone=LLaVA-7B2024.03 | 64.53 | — | |
| PVIT-7BEvaluation Protocol=Ground-truth boxes, Parameters=7B2024.06 | 64.5 | — | |
| GPT4ROIPT=266K, IT=731K, Vision Backbone=ViT-L, LLM Backbone=LLaVA-7B2024.03 | 64.01 | — | |
| GPT4RoI-7BEvaluation Protocol=Ground-truth boxes, Parameters=7B2024.06 | 64 | — | |
| GSCG ClassifierAblation=no neighbors/global context2025.12 | 56.66 | — | |
| ShikraPT=600K, IT=5.5M, Vision Backbone=ViT-L, LLM Backbone=Vicuna-7B2024.03 | 53.91 | — | |
| Shikra-7BEvaluation Protocol=Ground-truth boxes, Parameters=7B2024.06 | 53.9 | — | |
| ResNetModality=Combined 6-channel2025.12 | 53.52 | — | |
| ResNetContext Configuration=Object-Only, no context2025.12 | 44.1 | — | |
| Llama 4 ScoutModality=Multimodal, Evaluation Protocol=Zero-Shot2025.12 | 42.34 | — | |
| LLaVAPT=595K, IT=158K, Vision Backbone=ViT-L, LLM Backbone=Vicuna-7B2024.03 | 40.04 | — | |
| LLaVA-7BEvaluation Protocol=Ground-truth boxes, Parameters=7B2024.06 | 40 | — | |
| ResNetContext Configuration=Full Context-Only, no object2025.12 | 38.45 | — | |
| GSCG ClassifierAblation=minimal model, no object information, no geometry, only with labels of immediate neighbors2025.12 | 38.41 | — | |
| Llama 4 ScoutModality=Text-Only, Evaluation Protocol=Few-Shot2025.12 | 28.4 | — | |
| Llama 4 ScoutModality=Image-Only, Context Configuration=Context2025.12 | 24.93 | — | |
| Llama 4 ScoutModality=Text-Only, Evaluation Protocol=Zero-Shot2025.12 | 23.16 | — | |
| ASMPT=~22M, IT=~22M, Vision Backbone=ViT-L, LLM Backbone=Hasky-7B2024.03 | — | 69.3 | |
| ASM-7BEvaluation Protocol=Ground-truth boxes, Parameters=7B2024.06 | — | 69.3 | |
| CLIPVision Backbone=ViT-L2024.03 | — | 58.9 | |
| CLIPEvaluation Protocol=Ground-truth boxes2024.06 | — | 58.9 | |
| RegionCLIPVision Backbone=R50x42024.03 | — | 58.3 | |
| RegionCLIPEvaluation Protocol=Ground-truth boxes2024.06 | — | 58.3 |