Fine-Grained Open-Vocabulary Object Detection on FG-OVD
66.1AP (Hard)ObjEmbed-4B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ObjEmbed-4Bparameters=4B2026.02 | 66.1 | 74.1 | 77.2 | 76 | — | |
| ObjEmbed-2Bparameters=2B2026.02 | 65.7 | 74.3 | 77.3 | 76.2 | — | |
| Qwen3-VL-4Bparameters=4B2026.02 | 62 | 62.6 | 64 | 45.7 | — | |
| Qwen3-VL-2Bparameters=2B2026.02 | 59.3 | 58.9 | 53.8 | 55.9 | — | |
| GUIDED2026.02 | 57.5 | 69.5 | 73.3 | 72.6 | — | |
| FG-CLIPtrained_with_dense_annotations=true2026.02 | 46.1 | 66.6 | 68.7 | 83.4 | — | |
| FG-CLIPBackbone=ViT-B/162025.12 | 46.1 | 66.6 | 68.7 | 83.4 | 66.2 | |
| WeDetect-Ref-4Bparameters=4B2026.02 | 31.1 | 45.7 | 50.1 | 68.4 | — | |
| WeDetect-Ref-2Bparameters=2B2026.02 | 28.7 | 42.5 | 48.1 | 65.6 | — | |
| FineCLIPtrained_with_dense_annotations=true2026.02 | 26.8 | 49.8 | 50.4 | 71.9 | — | |
| OWLv2 (L/14)backbone=L/142026.02 | 25.4 | 41.2 | 42.8 | 63.2 | — | |
| MulCLIPBackbone=ViT-B/162025.12 | 19.24 | 40.73 | 47.27 | 68.63 | 43.97 | |
| GOALBackbone=ViT-B/162025.12 | 18.65 | 39.66 | 44.5 | 72.78 | 43.9 | |
| FineLIPBackbone=ViT-B/162025.12 | 18.17 | 38.88 | 41.96 | 73.79 | 43.2 | |
| Grounding-DINO-Tscale=Tiny2026.02 | 17 | 28.4 | 31 | 62.5 | — | |
| CAFTtrained_with_dense_annotations=false2026.02 | 16.8 | 32.8 | 38.6 | 62.5 | — | |
| LLMDet-Tscale=Tiny2026.02 | 15 | 26.2 | 23.8 | 55.4 | — | |
| EVA-CLIPtrained_with_dense_annotations=false2026.02 | 14 | 30.1 | 29.4 | 58.3 | — | |
| CLIPtrained_with_dense_annotations=false2026.02 | 12 | 23.1 | 22.2 | 58.5 | — | |
| Long-CLIPtrained_with_dense_annotations=false2026.02 | 9.2 | 18.4 | 16.2 | 51.8 | — |