Pedestrian crossing intention prediction on VR-based egocentric video dataset (test)
78.9AccuracyVLP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VLPGroup=Fine-tuned, Prompt=Standard, Time (s)=0.0072026.06 | 78.9 | 78.7 | |
| Qwen3-VL-2BGroup=Fine-tuned, Prompt=Standard, Time (s)=0.122026.06 | 77.7 | 77.3 | |
| CLIP+TransformerGroup=Baselines, Prompt=—2026.06 | 72.7 | 72.4 | |
| Qwen2.5-VL-7BGroup=Zero-shot, Prompt=Standard, Time (s)=0.33, Quantization=8-bit2026.06 | 63.2 | 57.9 | |
| Qwen2.5-VL-7BGroup=Zero-shot, Prompt=Visual prompt, Time (s)=0.33, Quantization=8-bit2026.06 | 60.5 | 52.6 | |
| Intern3VL-8BGroup=Zero-shot, Prompt=Standard, Time (s)=1.35, Quantization=8-bit2026.06 | 59.4 | 47.5 | |
| Qwen2.5-VL-7BGroup=Zero-shot, Prompt=CoT (multi), Time (s)=3.78, Quantization=8-bit2026.06 | 59.4 | 37.1 | |
| Qwen3-VL-2BGroup=Zero-shot, Prompt=Standard, Time (s)=0.122026.06 | 59.3 | 58 | |
| Intern3VL-2BGroup=Zero-shot, Prompt=Standard, Time (s)=0.722026.06 | 58.6 | 44.6 | |
| Qwen3-VL-8BGroup=Zero-shot, Prompt=Standard, Time (s)=0.32, Quantization=8-bit2026.06 | 57.9 | 55.9 | |
| Qwen2.5-VL-7BGroup=Zero-shot, Prompt=CoT (simple), Time (s)=2.32, Quantization=8-bit2026.06 | 56.8 | 54.3 | |
| MajorityGroup=Baselines, Prompt=—2026.06 | 56.7 | 36.2 | |
| VLPGroup=Zero-shot, Prompt=Standard, Time (s)=0.012026.06 | 56.7 | 37.8 | |
| RandomGroup=Baselines, Prompt=—2026.06 | 50 | 49.7 |