PPE-violation reasoning and localization on Construction Site Monitoring VI (val)
5.6Average MarkQwen2.5 (Baseline)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Qwen2.5 (Baseline)Evaluation protocol=LoRA-tuning, Model scale=7B2026.07 | 5.6 | 16.8 | 68.7 | 24.1 | 23.3 | 100 | |
| AVA-VLMEvaluation protocol=LoRA-tuning, Model scale=7B, Downsampling=1/4 downsampled global images2026.07 | 5.4 | 45.2 | 50.9 | 33.9 | 28.4 | 30.6 | |
| Qwen2.5⋆Evaluation protocol=LoRA-tuning, Model scale=7B2026.07 | 5.1 | 52.1 | 48.8 | 35.2 | 24.1 | 193 | |
| GPTEvaluation protocol=Few-shot, Number of shots=5-shot2026.07 | 4.7 | — | 14 | — | — | — | |
| GPTEvaluation protocol=Zero-shot / Few-shot2026.07 | 4.5 | — | 11.9 | — | — | — | |
| LLaVA-NeXTEvaluation protocol=Few-shot, Model scale=34B, Number of shots=1-shot2026.07 | 3.9 | — | — | — | — | — | |
| LLaVA-v1.5 CoTEvaluation protocol=Zero-shot / Few-shot, Model scale=13B, Chain-of-Thought (CoT)=true2026.07 | 3.1 | — | 6.9 | — | — | — | |
| LLaVA-v1.5Evaluation protocol=Zero-shot / Few-shot, Model scale=13B2026.07 | 2.9 | — | 20 | — | — | — | |
| Qwen2.5Evaluation protocol=Zero-shot, Model scale=7B2026.07 | 0 | 0 | 0 | 0 | 0 | 100 |