Aerial Object Navigation on AirSim v1.0 (All scenes)
0.805SRHuman Agent
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Human AgentParticipant Type=Professional drone operators, Training=>20 hours2026.01 | 0.805 | 0.817 | 0.617 | 18.8 | 24.6 | |
| AirHuntVLM=Qwen-VL-Max, Object Detection=Grounding-DINO, Segmentation=SAM22026.01 | 0.731 | 0.805 | 0.42 | 11.6 | 120.8 | |
| RecTask1Variant=Egocentric observations for Task 1 VLM queries2026.01 | 0.593 | 0.688 | 0.407 | 25.6 | 128.7 | |
| Two-stageVariant=Two-stage planning2026.01 | 0.589 | 0.634 | 0.325 | 21.4 | 198.2 | |
| ScalarizationVariant=Scalarization strategy2026.01 | 0.332 | 0.389 | 0.175 | 54.2 | 239.4 | |
| RecTask2Variant=Raw observations for visibility determination2026.01 | 0.312 | 0.384 | 0.125 | 65.1 | 135.6 | |
| PRPSearcherVLM=Qwen-VL-Max2026.01 | 0.24 | 0.266 | 0.193 | 90.2 | 324.6 | |
| FlySearchVLM=Qwen-VL-Max, Input=RGB images with coordinate grid2026.01 | 0.203 | 0.342 | 0.086 | 137.5 | 767.7 | |
| Value-greedyVariant=Greedy action selection2026.01 | 0.162 | 0.183 | 0.055 | 105.6 | 312.7 | |
| UAV-onVLM=Qwen-VL-Max, Input=RGB frames, text, depth, pose history2026.01 | 0.074 | 0.333 | 0.038 | 64.9 | 364.13 | |
| STARSearcherVLM=Qwen-VL-Max, Logic=Geometric-only hierarchical planner2026.01 | 0.024 | 0.049 | 0.011 | 103.5 | 296.7 |