ObjectGoal Navigation on MP3D (val)
64Success RateHydra-Nav-IRFT (Stage 3)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Hydra-Nav-IRFT (Stage 3)Observation=RGB + Depth, Action Output=Low + High, Dual Architecture=true2026.02 | 64 | 29.6 | — | |
| VLingNav2026.01 | 58.9 | 26.5 | — | |
| RIMMemory=Implicit2025.11 | 50.3 | 17 | — | |
| Hydra-Nav-SFT (Stage 2)Observation=RGB + Depth, Action Output=Low + High, Dual Architecture=true2026.02 | 49 | 16.5 | — | |
| DSCDZero-Shot=true, Training-Free=true, base VLM=Gemini-2.5-Flash-Lite2026.01 | 47.8 | 24.2 | — | |
| VLingNavtraining=Supervised Fine-Tuning (SFT)2026.01 | 47.4 | 25.8 | — | |
| CogNavObservation=RGB + Depth, Action Output=High, Dual Architecture=false2026.02 | 46.6 | 16.1 | — | |
| CogNavZero-Shot=true, Training-Free=true2026.01 | 46.6 | 16.1 | — | |
| CogNav2026.01 | 46.6 | 16.1 | — | |
| WMNavObservation=RGB + Depth, Action Output=High, Dual Architecture=false2026.02 | 45.2 | 17.2 | — | |
| CompassNavObservation=RGB + Depth, Action Output=High, Dual Architecture=false2026.02 | 42 | 17.5 | — | |
| UniGoalObservation=RGB + Depth, Action Output=High, Dual Architecture=false2026.02 | 41 | 16.4 | — | |
| UniGoal2026.01 | 41 | 16.4 | — | |
| SG-Nav-GPTUnsupervised=true, Zero-shot=true, LLM=GPT-4-0613, VLM=LLaVA-1.62024.10 | 40.2 | 16 | — | |
| SG-NavObservation=RGB + Depth, Action Output=High, Dual Architecture=false2026.02 | 40.2 | 16 | — | |
| SG-NavZero-Shot=true, Training-Free=true2026.01 | 40.2 | 16 | — | |
| SG-Nav2026.01 | 40.2 | 16 | — | |
| SG-Nav-LLAMAUnsupervised=true, Zero-shot=true, LLM=LLaMA-7B, VLM=LLaVA-1.62024.10 | 40.1 | 16 | — | |
| ApexNav2026.01 | 39.2 | 17.8 | — | |
| SGMZero-Shot=true, Training-Free=false2026.01 | 37.7 | 14.7 | — | |
| SGMMemory=Explicit2025.11 | 37.7 | 14.7 | — | |
| BeliefMapNavObservation=RGB + Depth, Action Output=High, Dual Architecture=false2026.02 | 37.3 | 17.6 | — | |
| OpenFMNavUnsupervised=true, Zero-shot=true2024.10 | 37.2 | 15.7 | — | |
| OpenFMNavZero-Shot=true, Training-Free=true2026.01 | 37.2 | 15.7 | — | |
| VLFMObservation=RGB + Depth, Action Output=High, Dual Architecture=false2026.02 | 36.4 | 17.5 | — | |
| VLFMZero-Shot=true, Training-Free=false2026.01 | 36.4 | 17.5 | — | |
| VLFM2026.01 | 36.4 | 17.5 | — | |
| VLFMMemory=Explicit2025.11 | 36.4 | 17.5 | — | |
| VLFM†Memory=Explicit, Segmentation=Proposed model from this paper2025.11 | 36.3 | 17.1 | — | |
| VLFMUnsupervised=true, Zero-shot=true2024.10 | 36.2 | 15.9 | — | |
| SemEXPUnsupervised=false, Zero-shot=false2024.10 | 36 | 14.4 | — | |
| SemEXPZero-Shot=false, Training-Free=false2026.01 | 36 | 14.4 | — | |
| Habitat-WebMemory=Implicit2025.11 | 35.4 | 10.2 | — | |
| TopV-NavZero-Shot=true, Training-Free=true2026.01 | 35.2 | 16.4 | — | |
| FOM-NavMemory=Explicit2025.11 | 35 | 18.7 | — | |
| L3MVNUnsupervised=true, Zero-shot=true2024.10 | 34.9 | 14.5 | — | |
| L3MVNObservation=RGB + Depth, Action Output=High, Dual Architecture=false2026.02 | 34.9 | 14.5 | — | |
| L3MVNZero-Shot=true, Training-Free=true2026.01 | 34.9 | 14.5 | — | |
| 6-Action AgentNumber of Actions=6, RedNet Fine-tuned=true, Tethered Training=false2021.04 | 34.6 | 7.93 | — | |
| Red-RabbitMethod Type=End-to-end RL2022.01 | 34.6 | 7.9 | — | |
| 4-Action AgentNumber of Actions=4, RedNet Fine-tuned=false, Tethered Training=false2021.04 | 34.4 | 9.58 | — | |
| VLFM*Memory=Explicit, Implementation=Reproduced2025.11 | 34.4 | 16.7 | — | |
| 4-Action AgentNumber of Actions=4, RedNet Fine-tuned=true, Tethered Training=false2021.04 | 33.1 | 6.89 | — | |
| PONIMethod Type=Modular2022.01 | 31.8 | 12.1 | 5.1 | |
| PONIUnsupervised=false, Zero-shot=false2024.10 | 31.8 | 12.1 | — | |
| PONIZero-Shot=false, Training-Free=false2026.01 | 31.8 | 12.1 | — | |
| Habitat-WebZero-Shot=true, Training-Free=false2026.01 | 31.6 | 8.5 | — | |
| Hydra-Nav-Base (Stage 1)Observation=RGB + Depth, Action Output=Low + High, Dual Architecture=true2026.02 | 30.9 | 14.6 | — | |
| 6-Action AgentNumber of Actions=6, RedNet Fine-tuned=false, Tethered Training=false2021.04 | 30.8 | 7.6 | — | |
| 6-Action AgentNumber of Actions=6, RedNet Fine-tuned=true, Tethered Training=true2021.04 | 30.3 | 10.8 | — | |
| Predict-xyMethod Type=Non-interactive2022.01 | 29.4 | 10.7 | 5.5 | |
| Predict-θMethod Type=Non-interactive2022.01 | 29 | 10.6 | 5.7 | |
| ESCUnsupervised=true, Zero-shot=true2024.10 | 28.7 | 14.2 | — | |
| ESCObservation=RGB + Depth, Action Output=Low + High, Dual Architecture=false2026.02 | 28.7 | 11.2 | — | |
| ESCZero-Shot=true, Training-Free=true2026.01 | 28.7 | 14.2 | — | |
| OVRL2026.01 | 28.6 | 7.4 | — | |
| THDAMethod Type=End-to-end RL2022.01 | 28.4 | 11 | 5.6 | |
| ANSMethod Type=Modular2022.01 | 27.3 | 9.2 | 5.8 | |
| 6-Action AgentNumber of Actions=6, RedNet Fine-tuned=false, Tethered Training=true2021.04 | 26.6 | 9.79 | — | |
| FBEMethod Type=Modular2022.01 | 22.7 | 7.2 | 6.7 | |
| ZSONUnsupervised=true, Zero-shot=false2024.10 | 15.3 | 4.8 | — | |
| zsonObservation=RGB, Action Output=Low, Dual Architecture=false2026.02 | 15.3 | 4.8 | — | |
| ZSONZero-Shot=false, Training-Free=false2026.01 | 15.3 | 4.8 | — | |
| DD-PPOMethod Type=End-to-end RL2022.01 | 8 | 1.8 | 6.9 | |
| CoWUnsupervised=true, Zero-shot=true2024.10 | 7.4 | 3.7 | — | |
| CoWZero-Shot=true, Training-Free=true2026.01 | 7.4 | 3.7 | — | |
| BCMethod Type=Non-interactive2022.01 | 3.8 | 2.1 | 7.5 | |
| Predict-AMethod Type=Non-interactive2022.01 | 2.7 | 1.6 | 7.8 |