Position Estimation on MMAUD
0.11Dx ErrorTAME
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TAMEYear=2023, Category (Training)=Supervised, Modal (Inference)=Audio-Only2026.03 | 0.11 | 0.3 | 0.34 | 0.55 | 29.72 | |
| AV-FDTIYear=2024, Category (Training)=Supervised, Modal (Inference)=Audio+Image2026.03 | 0.13 | 0.23 | 0.36 | 0.53 | 11.57 | |
| AAUTEYear=2025, Category (Training)=Self-Supervised, Modal (Inference)=Audio-Only2026.03 | 0.14 | 0.26 | 0.25 | 0.48 | 59.08 | |
| Lei et al.Year=2026, Category (Training)=Self-Supervised, Modal (Inference)=Audio+Image2026.03 | 0.15 | 0.16 | 0.19 | 0.29 | 17.8 | |
| OursCategory (Training)=Zero-shot, Modal (Inference)=Image-Only2026.03 | 0.17 | 0.15 | 0.44 | 0.3 | 23.5 | |
| YOLOv12Year=2025, Category (Training)=Supervised, Modal (Inference)=Image-Only2026.03 | 0.21 | 0.52 | 0.33 | 0.54 | 20.11 | |
| YOLOv10Year=2024, Category (Training)=Supervised, Modal (Inference)=Image-Only2026.03 | 0.23 | 0.43 | 0.46 | 0.72 | 18.34 | |
| YOLO26Year=2026, Category (Training)=Supervised, Modal (Inference)=Image-Only2026.03 | 0.24 | 0.37 | 0.43 | 0.59 | 24.43 | |
| AV-PEDYear=2023, Category (Training)=Self-Supervised, Modal (Inference)=Audio+Image2026.03 | 0.31 | 0.43 | 0.52 | 0.87 | 12.54 | |
| AV-DTECYear=2024, Category (Training)=Self-Supervised, Modal (Inference)=Audio+Image2026.03 | 0.33 | 0.25 | 0.27 | 0.58 | 13.66 |