Embodied Visual Tracking on EVT-Bench Single Target Tracking
88.4SRNavFoM*
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| NavFoM*2026.01 | 88.4 | 80.7 | — | |
| VLingNav2026.01 | 88.4 | 81.2 | 2.07 | |
| VLingNavmode=SFT2026.01 | 87.2 | 78.9 | 1.23 | |
| NavFoM2026.01 | 86 | 80.5 | — | |
| TrackVLA++2026.01 | 86 | 81 | 2.1 | |
| TrackVLA2026.01 | 85.1 | 78.6 | 1.65 | |
| Uni-NaVid2026.01 | 53.3 | 67.2 | 12.6 | |
| IBVSopen-vocabulary detector=GroundingDINO [32]2026.01 | 42.9 | 56.2 | 3.75 | |
| EVTvisual foundation model=SoM [61] with GPT-4o [36]2026.01 | 32.5 | 49.9 | 40.5 | |
| EVT2026.01 | 24.4 | 39.1 | 42.5 | |
| PoliFormeropen-vocabulary detector=GroundingDINO [32]2026.01 | 4.67 | 15.5 | 40.1 |