Storage Location Prediction on Real-World Evaluation Dataset 1.0 (test)
38AccuracyHuman Annotator 1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Human Annotator 1IoU Threshold=1.02025.12 | 38 | 0.38 | |
| Human Annotator 3IoU Threshold=1.02025.12 | 36 | 0.361 | |
| Human Annotator 2IoU Threshold=1.02025.12 | 27 | 0.271 | |
| NOAM LLaMA-3.3IoU Threshold=1.0, Base Model=LLaMA-3.32025.12 | 23 | 0.232 | |
| NOAM GPT-4IoU Threshold=1.0, Base Model=GPT-42025.12 | 23 | 0.23 | |
| Grounding-DINOIoU Threshold=>= 0.952025.12 | 17 | 0.188 | |
| Grounding-DINOIoU Threshold=1.02025.12 | 13 | 0.188 | |
| Grounding-DINOIoU Threshold=1.0, Prompt=no item in prompt2025.12 | 10 | 0.117 | |
| GPT-4o APIIoU Threshold=>= 0.52025.12 | 8 | 0.082 | |
| RandomIoU Threshold=1.02025.12 | 6 | 0.062 | |
| Qwen-2.5IoU Threshold=>= 0.5, Checkpoint=Qwen/Qwen2.5-VL-72B-Instruct2025.12 | 5 | 0.091 | |
| Kosmos-2IoU Threshold=>= 0.52025.12 | 4 | 0.042 | |
| Gemini-1.5-flashIoU Threshold=>= 0.52025.12 | 3 | 0.034 | |
| Gemini-2.5-flashIoU Threshold=>= 0.52025.12 | 1 | 0.027 | |
| LLaMA-4IoU Threshold=>= 0.5, Checkpoint=meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP82025.12 | 1 | 0.094 |