Personalized Object Detection on POD 5-shot
80.7AccuracyLLaVa (hi) + CLIP (zV,i)
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaVa (hi) + CLIP (zV,i)Model Category=Teacher Upper Bound (Frozen VLM features), Ground-Truth Boxes=No2025.09 | 80.7 | |
| LLaVa + CLIP + PCA (512)Model Category=Teacher Upper Bound (Frozen VLM features), Ground-Truth Boxes=No2025.09 | 80.1 | |
| CLIP (zV,i)Model Category=Teacher Upper Bound (Frozen VLM features), Ground-Truth Boxes=No2025.09 | 77.8 | |
| LLaVa (hi)Model Category=Teacher Upper Bound (Frozen VLM features), Ground-Truth Boxes=No2025.09 | 72.7 | |
| MOCHA (YOLOv8n)Model Category=Student (Efficient Edge Deployment), Ground-Truth Boxes=No2025.09 | 45.9 |