Vision-and-Language Navigation on R4R unseen (val)
47Success Rate (SR)Ours
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| OursMapping Technique=3D Gaussian Map2026.05 | 47 | — | 6.05 | — | 60 | 52 | 35 | — | |
| VLN↻Bert+REMsetting=single-run2021.06 | 46 | — | 6.21 | 38.1 | 44.9 | 46.3 | 22.7 | — | |
| VLN-SIG2023.04 | 45.8 | — | — | — | 59.1 | 52.9 | 33.6 | — | |
| HAMT2026.05 | 45 | — | 6.09 | — | 58 | 50 | 32 | — | |
| HAMT2023.04 | 44.6 | — | — | — | 57.7 | 50.3 | 31.8 | — | |
| HAMTvariant=encoder-decoder2021.10 | 44.6 | — | 6.09 | — | 57.7 | 50.3 | 31.8 | — | |
| RecBERT2026.05 | 44 | — | 6.67 | — | 51 | 45 | 30 | — | |
| RecBERT2021.10 | 43.6 | — | 6.67 | — | 51.4 | 45.1 | 29.9 | — | |
| VLN↻Bertsetting=single-run, reproduced=true2021.06 | 42.5 | — | 6.48 | 32.4 | 41.4 | 41.8 | 20.9 | — | |
| NvEM+SEvol2022.04 | 39 | — | 6.9 | 29 | 41 | 0.36 | 20 | — | |
| NvEM2022.04 | 38 | — | 6.85 | 28 | 41 | 0.36 | 20 | — | |
| EnvDrop+REMsetting=single-run2021.06 | 37.9 | — | 8.21 | 25 | 42.3 | 39.7 | 18.5 | — | |
| RelGraph2021.10 | 36 | — | 7.43 | — | 41 | 47 | 34 | — | |
| RelGraph2026.05 | 36 | — | 7.43 | — | 41 | 47 | 34 | — | |
| RelGraph2022.04 | 35 | — | 7.55 | 25 | 37 | 0.32 | 18 | — | |
| EnvDropsetting=single-run, reproduced=true2021.06 | 34.7 | — | 9.18 | 21 | 37.3 | 34.7 | 12.1 | — | |
| IL+RL+REMsetting=single-run2021.06 | 33.1 | — | 8.83 | 20.1 | 38.6 | 37.6 | 15.7 | — | |
| SSM2021.03 | 32 | 22.1 | 8.27 | — | 53 | 39 | 19 | — | |
| SSM2026.05 | 32 | — | 8.27 | — | 53 | 39 | 19 | — | |
| IL+RLsetting=single-run, reproduced=true2021.06 | 31.9 | — | 8.88 | 18.7 | 32.3 | 31.7 | 12.2 | — | |
| OAAM2021.03 | 31 | 13.8 | — | — | 40 | — | 11 | — | |
| OAAM2022.04 | 31 | — | — | 23 | 40 | — | 11 | — | |
| EGPsetting=single-run2021.06 | 30.2 | — | 8 | — | 44.4 | 37.4 | 17.5 | — | |
| EGP2021.10 | 30.2 | — | 8 | — | 44.4 | 37.4 | 17.5 | — | |
| EGP2021.03 | 30 | 18.3 | 8 | — | 44 | 37 | 18 | — | |
| EGP2026.05 | 30 | — | 8 | — | 44 | 37 | 18 | — | |
| BABYWALKpre-trained with data augmentation=false2020.05 | 29.6 | 23.8 | 7.9 | 14 | 47.8 | 38.1 | 18.1 | — | |
| RCMoptimization=goal oriented2021.03 | 29 | 32.5 | 8.45 | — | 20 | 22 | 11 | — | |
| E-Drop2021.03 | 29 | 27 | — | — | 34 | — | 9 | — | |
| EnvDrop2022.04 | 29 | — | — | 18 | 34 | — | 9 | — | |
| RCM-b2022.04 | 29 | — | — | 21 | 35 | 0.3 | 13 | — | |
| RCM2021.10 | 29 | — | — | — | 35 | 30 | 13 | — | |
| RCM2026.05 | 29 | — | — | — | 35 | 30 | 13 | — | |
| RCM(GOAL)pre-trained with data augmentation=true2020.05 | 28.7 | 12.3 | 7.9 | 22.1 | 36.3 | 31.3 | 13.2 | — | |
| BABYWALKpre-trained with data augmentation=true2020.05 | 27.3 | 19 | 8.2 | 14.7 | 49.4 | 39.6 | 17.3 | — | |
| BabyWalksetting=single-run2021.06 | 27.3 | — | 8.2 | 14.7 | 49.4 | 39.6 | 17.3 | — | |
| PTAlevel=low-level2021.03 | 27 | 10.2 | 8.19 | — | 35 | 20 | 8 | — | |
| RCMsetting=single-run2021.06 | 26.1 | — | 8.08 | 7.7 | 34.6 | — | — | — | |
| RCMoptimization=fidelity oriented2021.03 | 26 | 28.5 | 8.08 | — | 35 | 30 | 13 | — | |
| SEQ2SEQpre-trained with data augmentation=false2020.05 | 25.7 | 28.5 | 8.5 | 14.1 | 20.7 | 20.6 | 9 | — | |
| SFpre-trained with data augmentation=true2020.05 | 24.9 | 26.1 | 8.3 | 16 | 23.6 | 22.7 | 9.2 | — | |
| RCM(FIDELITY)pre-trained with data augmentation=true2020.05 | 24.7 | 26.4 | 8.4 | 11.6 | 39.2 | 31.3 | 13.7 | — | |
| Speaker-Follower2021.03 | 24 | 19.9 | 8.47 | — | 30 | — | — | — | |
| PTAlevel=high-level2021.03 | 24 | 17.7 | 8.25 | — | 37 | 32 | 10 | — | |
| PTAsetting=single-run2021.06 | 24 | — | 8.25 | 10 | 37 | 32 | 10 | — | |
| SF2021.10 | 24 | — | 8.47 | — | 30 | — | — | — | |
| PTA2021.10 | 24 | — | 8.25 | — | 37 | 32 | 10 | — | |
| SF2026.05 | 24 | — | 8.47 | — | 30 | — | — | — | |
| Speaker-Followersetting=single-run2021.06 | 23.8 | — | 8.47 | 12.2 | 29.6 | — | — | — | |
| BabyWalkTraining Dataset=R2R2020.05 | 22.5 | 19.5 | 8.9 | 12.6 | 50.3 | 38.9 | 14.5 | — | |
| BabyWalkTraining Dataset=R2R, Data Augmentation=true2020.05 | 21.4 | 17.9 | 8.9 | 11.9 | 51 | 40.3 | 13.8 | — | |
| DACoZero-shot=true, Backbone=Qwen3-VL-8B2026.02 | 20.5 | — | 8.91 | 17.4 | — | — | — | 55 | |
| RegretfulTraining Dataset=R2R, Data Augmentation=true, Reimplemented or readapted from original authors=true2020.05 | 19.2 | 15.5 | 8.4 | 10.1 | 46.4 | 31.6 | 9.8 | — | |
| SEQ2SEQTraining Dataset=R2R2020.05 | 18.3 | 28.6 | 9.1 | 7.9 | 29.8 | 25.1 | 7.1 | — | |
| Speaker-Follower (SF)Training Dataset=R2R, Data Augmentation=true2020.05 | 16.7 | 28.9 | 9 | 7.4 | 30 | 25.3 | 6.7 | — | |
| RCM (Fidelity)Training Dataset=R2R, Data Augmentation=true2020.05 | 15.2 | 14.1 | 9.3 | 8.9 | 41.2 | 32.4 | 7.2 | — | |
| MapGPTZero-shot=true, Backbone=Qwen3-VL-8B2026.02 | 15.1 | — | 8.94 | 13 | — | — | — | 54.8 | |
| NavGPTZero-shot=true, Backbone=Qwen3-VL-8B2026.02 | 15 | — | 9.9 | 14.1 | — | — | — | 44 | |
| RCM (Goal)Training Dataset=R2R, Data Augmentation=true2020.05 | 14.7 | 13.2 | 9.2 | 8.9 | 42.5 | 33.3 | 7.3 | — | |
| FASTTraining Dataset=R2R, Data Augmentation=true, Reimplemented or readapted from original authors=true2020.05 | 13.3 | 29.7 | 9.1 | 7.7 | 41.8 | 33.5 | 7.2 | — |